Median forecast: MMLU state of the art 77.75% by end of 2024 (actual 88.7%)
“40. “Massive Multitask Language Understanding” Benchmark 77.75% 88.7% 1.59 32”
Context: Existential Risk Persuasion Tournament (XPT), June-Oct 2022; group median forecasts as reported by the Forecasting Research Institute in "Assessing Near-Term Accuracy in the XPT" (2025). Table row: question, median forecast, resolution, standardized error, N.
How we judge it
Our interpretation: treated as a point forecast of the best MMLU score by end-2024; correct if within 5 points of the resolved value.
Resolution criteria are our interpretation of the claim, written before grading. Methodology
- Said
- Resolve by
- Verdict
- Wrong ()
- Stated probability
- None stated
- Specificity
- 3 of 3 (precise and measurable)
- Score
- 0.0 points × weight 3
Evidence & grader’s note
Superforecasters' median of 77.75% undershot the resolved 88.7% by 11 points, one of their ten most surprising misses.