- Score
- —
- 95% interval
- —
- Resolved
- 0
- Record (C–P–W)
- 0–0–0
- Open
- 4
- Too vague
- 0
Not ranked yet: leaderboards need at least 5 resolved claims (has 0). Score = specificity-weighted average of claim points; we grade claims, not people.
Calibration
No resolved claims with a stated probability yet, so there is nothing to calibrate. Calibration needs forecasts like “70% chance of X”.
Track record
Open (4)
~25%: 80%-reliability one-month task horizons by the start of 2028
“We're pretty unlikely (25%?) to see 80% reliability at 1 month tasks by the start of 2028.”
METR 50% time horizons of about two weeks around the start of 2028
“I expect doubling times of around 170 days on METR's task suite (or similar tasks) [3] over the next 2 years or so which implies we'll be hitting 2 week 50% reliability horizon len…”
45%: AI capable of fully automating AI R&D by the start of 2033
“I now would put ~15% probability by the start of 2029 and a 45% chance by the start of 2033.”
~15%: AI capable of fully automating AI R&D by the start of 2029
“I now would put ~15% probability by the start of 2029 and a 45% chance by the start of 2033.”
Embed this scorecard
Right of reply: send a correction or response.