Mid-2025: AI agents impressive in theory but unreliable in practice
“The agents are impressive in theory (and in cherry-picked examples), but in practice unreliable.”
Context: AI 2027 scenario (Apr 3, 2025), which the authors called their "best guess" and later clarified had 2027 as the modal year, with somewhat longer medians. Section "Mid 2025: Stumbling Agents".
How we judge it
Our interpretation: correct if in mid/late 2025 general-purpose agents are widely reported as unreliable outside cherry-picked demos.
Resolution criteria are our interpretation of the claim, written before grading. Methodology
- Said
- Resolve by
- Verdict
- Correct ()
- Stated probability
- None stated
- Specificity
- 1 of 3 (vague or hedged)
- Score
- 1.0 points × weight 1
Evidence & grader’s note
Evidence timeline
Dated entries from the claim’s own record (statement, verdict, later developments) and sourced entries added by editors. Each links to its source; none changes the verdict.
- Scenario step
- AI 2027: Mid 2025, Stumbling Agents
Share this scorecard
Post on X LinkedIn Share image (PNG)
Embed / Cite
The embedded card is live: it shows the current verdict, with the date the data was read. If a verdict is corrected later, the card changes too, so cite the page and its date for a fixed record.