Methodology

This page explains exactly how a quote becomes a verdict and how verdicts become scores. If you think we have applied it wrongly to a claim, use Corrections & right of reply.

  • 209 claims tracked from 83 claimants
  • 84 resolved and scored; 121 open (5 past due, awaiting a verdict)
  • 3 listed as too vague to grade (1% of all claims)
  • 19 resolved claims carried a stated probability (mean Brier 0.291)

What we track

  • Predictions: public statements about what will happen (e.g. “AGI by 2027”, “robotaxis in 10 cities next year”).
  • Promises: public commitments an organization or official makes about its own actions (e.g. “we will release an open-weight model this summer”).
  • Only on-the-record statements by public figures, companies and public bodies. We do not track private individuals.

Sourcing rules

  • Every quote is verbatim from a primary source (the speaker’s own blog, transcript, filing or video) or a reputable outlet that quotes them directly.
  • We save the page text when we add a claim, and our build checks that the quote appears word-for-word in that saved copy. Quotes that cannot be verified are dropped.
  • We link an archive.org snapshot where one is available, so the source survives edits and deletions.
  • The “date said” is the date of the statement. If we only know the month or year, we show only the month or year.

Resolution criteria (our interpretation)

Many forecasts are loose. Before grading, we write down how we will judge the claim: what counts, by when, and from which evidence. These criteria are our interpretation, are shown on every claim page, and can be challenged. We read claims charitably but literally: if someone says “by the end of the year”, we use 31 December of that year.

Verdicts

  • Correct / Kept: the criteria were met by the deadline.
  • Partly correct / Partly kept: a material part came true, or it came true but late or in a narrower form than stated.
  • Wrong / Broken: the criteria were not met by the deadline.
  • Too vague to grade: no reasonable criteria could settle it. These are listed but never scored, and the share of vague claims is shown as a statistic.
  • Withdrawn: the claimant publicly retracted the claim before the deadline. Not scored.
  • Open: the deadline has not passed, or an editor has not yet graded it.

Points for a single claim

  • If the claimant stated a probability (e.g. “70% chance”), we use the Brier score: (probability − outcome)², where the outcome is 1 if the event happened and 0 if not. Lower Brier is better (0 is perfect; always saying 50% earns 0.25). For the overall score we convert it to points: points = 1 − Brier.
  • If no probability was stated: correct = 1 point, partly correct = 0.5, wrong = 0.

Specificity weighting

Each claim gets a specificity rating from 1 to 3: 1 vague or hedged, 2 dated but needing some interpretation, 3 precise and measurable. A claimant’s score is the specificity-weighted average of their claim points, so bold, checkable calls count more than hedged ones, in both directions.

Leaderboards, sample size and uncertainty

  • Only resolved, scored claims count. Open, too-vague and withdrawn claims never affect the score.
  • A claimant is ranked only with at least 5 resolved claims. Fewer than that, and they appear as “not yet ranked”.
  • Every ranking shows n (the number of resolved claims) and a 95% Wilson interval on the unweighted hit rate (partly correct counts as half a success). Overlapping intervals mean we cannot really tell two claimants apart.
  • Ties are broken by the lower end of the interval, then by n.
  • Mean Brier is shown separately for claimants who stated probabilities, with its own count.

Calibration

For forecasts with a stated probability, profile pages show a calibration chart: forecasts are grouped into five probability bands, and we plot the average stated probability against how often those events actually happened. A well-calibrated forecaster’s points sit near the diagonal. With few probabilistic forecasts, the chart is mostly illustrative.

Promise tracker

Promises are graded the same way but labelled kept / partly kept / broken / pending. The kept rate is (kept + ½ partly) ÷ judged promises; pending promises are excluded.

Review and corrections

  • Every claim is checked by a human editor before it is published. Claims awaiting review are marked “pending editorial review” and are not shown in search results.
  • Open claims are re-checked monthly. When a deadline passes, an editor grades it using linked evidence and writes a grader’s note.
  • Anyone, especially the person quoted, can send a correction, reply or new evidence. Material changes are recorded in the grader’s note with the date.

Limits

Which claims get tracked is an editorial choice, and famous people make more quotable predictions. Scores describe a sample of public statements, not a person’s overall judgement or character. Read every score with its n and interval.

Data is available as JSON, an RSS feed of resolutions, and an llms.txt summary.