Organization
XPT superforecasters
Superforecaster panel in the Forecasting Research Institute's 2022 Existential Risk Persuasion Tournament (group median forecasts)
- Score
- 13%
- 95% interval
- 2%–53%
- Resolved
- 6
- Record (C–P–W)
- 0–0–6
- Open
- 0
- Too vague
- 0
- Mean Brier
- 0.620 (2)
Ranked on leaderboards (6 resolved claims). Score = specificity-weighted average of claim points; we grade claims, not people.
Calibration
| Stated probability | Forecasts | Mean stated | Happened |
|---|---|---|---|
| 0–20% | 1 | 2% | 100% |
| 20–40% | 0 | — | — |
| 40–60% | 1 | 54% | 0% |
| 60–80% | 0 | — | — |
| 80–100% | 0 | — | — |
Track record
Resolved (6)
53.5% that Robin Hanson wins his bet that GPT revenue stays under $1B
“35. GPT Revenue (Hanson Wins Bet that GPT Revenue < $1B) 53.5% 0% 1.07 32”
Median forecast: largest ML model would have 100 trillion parameters by 2024 (actual ~10 trillion)
“49. Largest Number of Parameters in a Machine Learning Model 100 trillion 10 trillion 1.71 31”
Median forecast for the maximum compute used in an AI experiment by 2024 was about 6x too low
“45. Maximum Compute Used in an AI Experiment 100,000 578,703.7 1.92 33”
Median forecast: MATH benchmark state of the art 71% by end of 2024 (actual 87.9%)
“39. MATH Dataset Benchmark 71% 87.92% 1.38 30”
Median forecast: MMLU state of the art 77.75% by end of 2024 (actual 88.7%)
“40. “Massive Multitask Language Understanding” Benchmark 77.75% 88.7% 1.59 32”
2.3%: AI gold-level performance at the IMO by 2025
“an outcome to which domain experts assigned only an 8.6% probability and superforecasters a mere 2.3% probability.”
Embed this scorecard
Right of reply: send a correction or response.