PhD-level AI for specific tasks in "a year and a half"
“And then in the next couple of years, we're looking at PhD-level intelligence for specific tasks.”
Context: Mira Murati (then OpenAI CTO) at Dartmouth Engineering's "AI Everywhere" conversation (June 2024; video posted 19 Jun 2024, uploaded captions). Asked "Meaning like a year from now?", she answered "Yeah, a year and a half let's say."
How we judge it
Our interpretation: by Dec 2025, a widely available AI system performs at PhD level on specific expert tasks (e.g. matches PhD experts on GPQA-style questions in their field).
Resolution criteria are our interpretation of the claim, written before grading. Methodology
- Said
- Resolve by
- Verdict
- Correct ()
- Stated probability
- None stated
- Specificity
- 2 of 3 (dated, some interpretation)
- Score
- 1.0 points × weight 2
Evidence & grader’s note
- OpenAI (Sep 12, 2024): o1 "surpassed the performance of those human experts" (PhDs) on GPQA Diamond
- OpenAI: Introducing GPT-5 (Aug 2025): "PhD-level intelligence"
Within the window, OpenAI's o1 (Sep 2024) became the first model to beat PhD experts recruited by OpenAI on GPQA Diamond, the example test in our criteria; OpenAI notes this does not make it "more capable than a PhD in all respects". GPT-5 (Aug 2025) was marketed as "PhD-level intelligence".
Evidence timeline
Dated entries from the claim’s own record (statement, verdict, later developments) and sourced entries added by editors. Each links to its source; none changes the verdict.
- Readiness milestone
- Part of AI matches top human experts on hard tests (says it happens)
Share this scorecard
Post on X LinkedIn Share image (PNG)
Embed / Cite
The embedded card is live: it shows the current verdict, with the date the data was read. If a verdict is corrected later, the card changes too, so cite the page and its date for a fixed record.