WrongPrediction

8% that AI gets IMO gold by 2025

“Maybe I'll go 8% on "gets gold" instead of "solves hardest problem."”

Paul ChristianoAI alignment researcherSaid · Source: lesswrong.com · Archived copy

Context: LessWrong post "IMO challenge bet with Eliezer" (final predictions from a Nov 2021 exchange).

How we judge it

An AI built before the IMO achieves a gold-medal score at the 2022-2025 IMO.

Resolution criteria are our interpretation of the claim, written before grading. Methodology

Said
Resolve by
Verdict
Wrong ()
Stated probability
8%
Specificity
3 of 3 (precise and measurable)
Score
Brier 0.846 → 0.15 points

Evidence & grader’s note

The bet covered an AI built before the IMO getting gold "under time controls etc. of grand challenge". The IMO Grand Challenge's proposed rules also asked for formal (Lean) proofs and an AI released as open source before the contest. At IMO 2025, Google DeepMind's Gemini Deep Think scored 35/42, graded by IMO coordinators, working in natural language within the 4.5-hour time limits. OpenAI reported 35/42 for an experimental model under the same time limits, graded by three former medalists. Both runs wrote natural-language proofs, not formal ones, but Christiano agreed that the related prediction market should resolve yes. Brier = (0.08 – 1)^2 = 0.846.

Share this scorecard

Post on X LinkedIn Share image (PNG)