8% that AI gets IMO gold by 2025
“Maybe I'll go 8% on "gets gold" instead of "solves hardest problem."”
Context: LessWrong post "IMO challenge bet with Eliezer" (final predictions from a Nov 2021 exchange).
How we judge it
An AI built before the IMO achieves a gold-medal score at the 2022-2025 IMO.
Resolution criteria are our interpretation of the claim, written before grading. Methodology
- Said
- Resolve by
- Verdict
- Wrong ()
- Stated probability
- 8%
- Specificity
- 3 of 3 (precise and measurable)
- Score
- Brier 0.846 → 0.15 points
Evidence & grader’s note
- Google DeepMind: officially gold-medal standard at IMO 2025
- Alexander Wei (OpenAI) on X: evaluated under contest rules, two 4.5-hour sessions, no tools (Jul 19, 2025)
- Alexander Wei (OpenAI) on X: 35/42, graded by three former IMO medalists (Jul 19, 2025)
- IMO Grand Challenge: proposed rules (formal proofs, time limits, open source)
- Manifold market on the bet: Christiano says "I think it should resolve yes"
- LessWrong: Google and OpenAI get 2025 IMO gold
The bet covered an AI built before the IMO getting gold "under time controls etc. of grand challenge". The IMO Grand Challenge's proposed rules also asked for formal (Lean) proofs and an AI released as open source before the contest. At IMO 2025, Google DeepMind's Gemini Deep Think scored 35/42, graded by IMO coordinators, working in natural language within the 4.5-hour time limits. OpenAI reported 35/42 for an experimental model under the same time limits, graded by three former medalists. Both runs wrote natural-language proofs, not formal ones, but Christiano agreed that the related prediction market should resolve yes. Brier = (0.08 – 1)^2 = 0.846.