WrongPrediction

2.3%: AI gold-level performance at the IMO by 2025

“an outcome to which domain experts assigned only an 8.6% probability and superforecasters a mere 2.3% probability.”

XPT superforecastersSuperforecaster panel in the Forecasting Research Institute's 2022 Existential Risk Persuasion Tournament (group median forecasts)Said · Source: forecastingresearch.org

Context: Existential Risk Persuasion Tournament (XPT), June-Oct 2022; group median forecasts as reported by the Forecasting Research Institute in "Assessing Near-Term Accuracy in the XPT" (2025).

How we judge it

Our interpretation: YES if an AI system achieves gold-medal-level performance at the International Mathematical Olympiad by 2025.

Resolution criteria are our interpretation of the claim, written before grading. Methodology

Said
Resolve by
Verdict
Wrong ()
Stated probability
2%
Specificity
3 of 3 (precise and measurable)
Score
Brier 0.955 → 0.05 points

Evidence & grader’s note

AI systems from Google DeepMind and OpenAI reached gold-medal level at IMO 2025. At p=0.023 on an event that happened, the Brier-based score is about 0.05.

Share this scorecard

Post on X LinkedIn Share image (PNG)