WrongPrediction

Median forecast: MMLU state of the art 77.75% by end of 2024 (actual 88.7%)

“40. “Massive Multitask Language Understanding” Benchmark 77.75% 88.7% 1.59 32”

XPT superforecastersSuperforecaster panel in the Forecasting Research Institute's 2022 Existential Risk Persuasion Tournament (group median forecasts)Said · Source: forecastingresearch.org

Context: Existential Risk Persuasion Tournament (XPT), June-Oct 2022; group median forecasts as reported by the Forecasting Research Institute in "Assessing Near-Term Accuracy in the XPT" (2025). Table row: question, median forecast, resolution, standardized error, N.

How we judge it

Our interpretation: treated as a point forecast of the best MMLU score by end-2024; correct if within 5 points of the resolved value.

Resolution criteria are our interpretation of the claim, written before grading. Methodology

Said
Resolve by
Verdict
Wrong ()
Stated probability
None stated
Specificity
3 of 3 (precise and measurable)
Score
0.0 points × weight 3

Evidence & grader’s note

Superforecasters' median of 77.75% undershot the resolved 88.7% by 11 points, one of their ten most surprising misses.

Share this scorecard

Post on X LinkedIn Share image (PNG)