Testing ocean “emulators” with a mid-Holocene climate experiment reveals what they get right — and what they miss
Scientists tested AI-style ocean models called emulators to see if they can predict how the ocean responds when the climate changes. An emulator is a fast model that learns from detailed climate simulations and then predicts how the ocean will evolve over months to centuries. The team used data from a midHolocene experiment of a numerical climate model as a deliberate “out-of-sample” test — a climate that was different from the one the emulators were trained on — to check whether the emulators can still capture the ocean’s forced response to surface changes.
The researchers focused on autoregressive, full-depth ocean emulators. “Autoregressive” means the emulator predicts the next time step using recent past information. They asked whether these emulators reproduce the ocean’s response to changed surface conditions, including different orbital forcings that alter seasonal sunlight. They compared the emulator output to the numerical model’s midHolocene experiment to evaluate spatial patterns, seasonal changes, and changes in ocean variability.
Overall, the emulators did well at finding the large-scale pattern of the forced response and at reproducing seasonal shifts and changes in the spatial structure of variability. However, they tended to underestimate how strong those changes were. Simpler baseline methods that infer the ocean state directly from surface boundary conditions could also recover much of the large-scale pattern, but only near the surface. Those baselines did not capture seasonal changes or changes in variability, which suggests that some representation of ocean dynamics is needed for those aspects.
A major limitation is that the emulators failed to reproduce the slow, internally driven evolution of the deep ocean. In other words, they struggled to capture long, internal changes that are not directly forced by the surface. The team also tracked training progress across epochs and found that matching average conditions in the training climate does not guarantee that an emulator has learned the dynamics needed to respond correctly in a different climate.