2.3%: AI gold-level performance at the IMO by 2025
“an outcome to which domain experts assigned only an 8.6% probability and superforecasters a mere 2.3% probability.”
Context: Existential Risk Persuasion Tournament (XPT), June-Oct 2022; group median forecasts as reported by the Forecasting Research Institute in "Assessing Near-Term Accuracy in the XPT" (2025).
How we judge it
Our interpretation: YES if an AI system achieves gold-medal-level performance at the International Mathematical Olympiad by 2025.
Resolution criteria are our interpretation of the claim, written before grading. Methodology
- Said
- Resolve by
- Verdict
- Wrong ()
- Stated probability
- 2%
- Specificity
- 3 of 3 (precise and measurable)
- Score
- Brier 0.955 → 0.05 points
Evidence & grader’s note
AI systems from Google DeepMind and OpenAI reached gold-medal level at IMO 2025. At p=0.023 on an event that happened, the Brier-based score is about 0.05.