LLM hallucinations "much less of a problem, certainly by 2029"
“LLM hallucinations [where they create nonsensical or inaccurate outputs] will become much less of a problem, certainly by 2029”
Context: Interview with The Guardian (Jun 2024). Bracketed gloss is the Guardian's.
How we judge it
Our interpretation: by end-2029, frontier models' hallucination rates on standard factuality benchmarks (e.g. SimpleQA-style or OpenAI/Vectara hallucination leaderboards) are a small fraction (≤ one-fifth) of 2024 levels, and hallucination is no longer named a top barrier to deployment.
Resolution criteria are our interpretation of the claim, written before grading. Methodology
- Said
- Resolve by
- Due in 1179 days
- Verdict
- Open
- Stated probability
- None stated
- Specificity
- 2 of 3 (dated, some interpretation)