# SiPoly > SiPoly is a public scoreboard of AI predictions and promises: verbatim quotes with primary sources, operationalized resolution criteria (our interpretation), and verdicts graded when the resolve-by date arrives. Claims are graded, not people. Tracking 276 claims from 102 claimants: 111 resolved, 159 open, 3 too vague to grade, 3 withdrawn. Data as of 2026-10-09. ## Key pages - [Methodology: scoring (Brier, specificity weighting, Wilson intervals)](https://sipoly.si/methodology/) - [Leaderboards (min. 5 resolved claims)](https://sipoly.si/leaderboards/) - [Coming-due calendar](https://sipoly.si/coming-due/) - [Promise tracker for organizations](https://sipoly.si/promises/) - [Editorial standards and disclosure](https://sipoly.si/editorial-standards/) - [Corrections and right of reply](https://sipoly.si/corrections/) ## Data - [API documentation](https://sipoly.si/api/): every endpoint, its parameters and examples - [JSON API: claims](https://sipoly.si/wp-json/sipoly/v1/claims) - [JSON API: leaderboard](https://sipoly.si/wp-json/sipoly/v1/leaderboard) - [JSON API: claimants](https://sipoly.si/wp-json/sipoly/v1/claimants) (paged: ?per_page=&page=) - [RSS: resolutions](https://sipoly.si/feed/resolutions/) - [Full claim list for LLMs](https://sipoly.si/llms-full.txt) ## Recently resolved - [Kept: An automated "intern-level research assistant" by September 2026](https://sipoly.si/claim/openai-research-intern-sep-2026/) — OpenAI - [Wrong: "In one year, the vast majority of programmers will be replaced by AI programmers"](https://sipoly.si/claim/schmidt-programmers-replaced-1y/) — Eric Schmidt - [Broken: More personalized Siri "in the coming year"](https://sipoly.si/claim/apple-personal-siri-coming-year/) — Apple - [Wrong: AI smarter than any one human "around the end of next year"](https://sipoly.si/claim/musk-ai-smarter-end-2025/) — Elon Musk - [Kept: China's core AI industry to exceed 400 billion RMB by 2025](https://sipoly.si/claim/china-aidp-core-industry-2025/) — China State Council - [Broken: Apollo Go robotaxis in 65 cities by 2025 (100 by 2030)](https://sipoly.si/claim/baidu-apollo-65-cities-2025/) — Baidu - [Broken: "One billion agents with Agentforce by the end of 2025"](https://sipoly.si/claim/salesforce-billion-agents-2025/) — Salesforce - [Correct: Models will work autonomously for full 8-hour days by mid-2026](https://sipoly.si/claim/schrittwieser-full-day-autonomy-mid-2026/) — Julian Schrittwieser - [Wrong: At least six Indian developers will build foundation models within 8-10 months](https://sipoly.si/claim/vaishnaw-india-foundation-models-8-10m/) — Ashwini Vaishnaw - [Correct: Mid-2025: specialized coding and research agents begin to transform their professions](https://sipoly.si/claim/ai2027-mid2025-coding-agents-transform/) — AI Futures Project ## All claims ### US real GDP growth in 2027 above 3.3%, ">50% higher than 2026" - Claimant: Elon Musk - Said: Sep 30, 2026 - Quote: "My guess for real GDP growth next year is >50% higher than 2026, so >3.3%" - Source: https://x.com/elonmusk/status/2105361313392476174 - Resolution criteria (our interpretation): Our interpretation: correct if the US Bureau of Economic Analysis's first estimate of real GDP growth for calendar 2027 (annual-average basis, normally published in late January 2028) is above 3.3%. - Resolve by: 2028-01-31 - Verdict: Open - URL: https://sipoly.si/claim/musk-us-gdp-2027-above-3-3pct/ ### "Some form of AGI in the next six to 18 months" (Sep 2026) - Claimant: Dan Schulman - Said: Sep 30, 2026 - Quote: "I think we get to some form of AGI in the next six to 18 months. I think it’s right on us." - Source: https://fortune.com/2026/09/30/this-revolution-is-happening-in-years-not-decades-verizon-ceo-dan-schulman-sees-agi-within-18-months-and-a-century-of-progress-in-a-decade/ - Resolution criteria (our interpretation): Our interpretation: correct if by 30 Mar 2028 (18 months after the statement) an AI system is broadly recognized by independent experts as AGI, i.e. able to do most economically valuable cognitive work at the level of a skilled professional across domains. A developer's own AGI claim does not count on its own. - Resolve by: 2028-03-30 - Verdict: Open - URL: https://sipoly.si/claim/schulman-agi-6-18-months/ ### About a 50% chance that AI companies automate the whole AI research process by end of 2028 - Claimant: Daniel Kokotajlo - Said: Sep 30, 2026 - Quote: "They are already starting to work on training AIs to do the whole research process, not just the coding. It’s unclear when they will succeed but my team and I think it could happen any year now. I personally would guess about 50% chance by end of 2028." - Source: https://blog.aifutures.org/p/senate-testimony-sept-2026 - Resolution criteria (our interpretation): Our interpretation: YES if by 31 Dec 2028 a frontier AI company has automated essentially the whole AI research and development process (AI systems, not human researchers, generate, run and interpret most of the research behind new frontier models, with humans mainly overseeing), as credibly reported by the company or independent evaluators. Stated probability 50%, so it is scored on calibration. - Resolve by: 2028-12-31 - Verdict: Open - URL: https://sipoly.si/claim/kokotajlo-ai-rnd-automation-50pct-2028/ ### EO 14434: proposed legal definition of "Super Intelligence" due to the President within 60 days - Claimant: The White House (U.S. executive branch) - Said: Sep 29, 2026 - Quote: "Within 60 days of the date of this order, the Assistant to the President for Science and Technology (APST), in consultation with the heads of other agencies as the APST deems appropriate, shall submit to the President proposed legislative language to establish a Federal definition of “Super Intelligence” and “SI”" - Source: https://www.whitehouse.gov/presidential-actions/2026/09/inaugurating-the-era-of-super-intelligence/ - Resolution criteria (our interpretation): Our interpretation: correct if by 28 Nov 2026 the APST has submitted the proposed legislative language for a federal definition of "Super Intelligence", as shown by a White House or OSTP statement, a transmitted bill text, or credible reporting; wrong if credible reporting shows it was not delivered on time. - Resolve by: 2026-11-28 - Verdict: Pending - URL: https://sipoly.si/claim/eo14434-si-definition-60d/ ### OpenAI's in-house chip program to start coming online in the first half of 2027 - Claimant: OpenAI - Said: Sep 29, 2026 - Quote: "which will start to come online in the first half of next of 2027 and will scale a lot in future years" - Source: https://www.youtube.com/watch?v=efcG5Uf-GH0 - Resolution criteria (our interpretation): Our interpretation: correct if by 30 Jun 2027 OpenAI's own custom AI chips are running production workloads (e.g. serving ChatGPT or API inference), as confirmed by OpenAI or reputable reporting. Tape-outs or test deployments alone do not count. - Resolve by: 2027-06-30 - Verdict: Pending - URL: https://sipoly.si/claim/openai-custom-chips-h1-2027/ ### Cybercab operating commercially in California "probably by the middle of next year" (2027) - Claimant: Tesla - Said: Sep 26, 2026 - Quote: "And, probably, by the middle of next year, it will be operating commercially in California." - Source: https://www.businessinsider.com/tesla-cybercab-timeline-mkbhd-hair-2026-9 - Resolution criteria (our interpretation): Our interpretation: correct if by 30 Jun 2027 Tesla's Cybercab (the two-seat vehicle with no steering wheel or pedals) carries paying members of the public in commercial service somewhere in California. - Resolve by: 2027-06-30 - Verdict: Pending - URL: https://sipoly.si/claim/tesla-cybercab-california-mid-2027/ ### SpaceX's AI effort to reach "pole position" in about six months (from Sep 2026) - Claimant: Elon Musk - Said: Sep 24, 2026 - Quote: "If our second derivative remains strong, SpaceX will reach pole position in about 6 months." - Source: https://x.com/elonmusk/status/2103160462472892536 - Resolution criteria (our interpretation): Our interpretation: correct if by 31 Mar 2027 a SpaceX/xAI model ranks #1 overall on the Artificial Analysis Intelligence Index or holds the top overall text score on LMArena (leading the frontier overall, not just a sub-index). - Resolve by: 2027-03-31 - Verdict: Open - URL: https://sipoly.si/claim/musk-spacex-ai-pole-position-6mo/ ### Embedded third-party evaluators with "employee-like access" inside Anthropic - Claimant: Anthropic - Said: Sep 23, 2026 - Quote: "First, we committed to embed external evaluators inside Anthropic with employee-like access, similar to a food inspector, and we recommended that other companies across the world do the same." - Source: https://transcripts.un.org/en/sc/10228.txt - Resolution criteria (our interpretation): Our interpretation: correct if by 31 Mar 2027 Anthropic and at least one named independent evaluator (e.g. METR) publicly confirm that an embedded third-party evaluation team has ongoing, employee-like access inside Anthropic (to models and training pipelines); wrong if no such arrangement is confirmed by then. - Resolve by: 2027-03-31 - Verdict: Pending - URL: https://sipoly.si/claim/anthropic-embedded-evaluators-2026/ ### A "country of geniuses in the data center" within "one or two years, maybe less" (Sep 2026) - Claimant: Dario Amodei - Said: Sep 23, 2026 - Quote: "The trajectory only needs to continue for a tiny period longer, one or two years, maybe less, to reach what I've called a country of geniuses in the data center." - Source: https://transcripts.un.org/en/sc/10228.txt - Resolution criteria (our interpretation): Our interpretation: correct if by 23 Sep 2028 independent evaluators (not only Anthropic) judge that AI systems meet Amodei's own description: models smarter than top human experts across most fields, able to work autonomously on tasks lasting days or weeks, and run as many parallel copies. - Resolve by: 2028-09-23 - Verdict: Open - URL: https://sipoly.si/claim/amodei-country-of-geniuses-1-2y/ ### Median 2029 for full coding automation (as of Feb 2026) - Claimant: Daniel Kokotajlo - Said: Feb 12, 2026 - Quote: "Mid-2028 is earlier than Daniel’s current median prediction for full coding automation (2029), but the 2-year takeoff to superintelligence is slower than his median takeoff speed of ~1 year." - Source: https://blog.aifutures.org/p/grading-ai-2027s-2025-predictions - Resolution criteria (our interpretation): Our interpretation: YES if by end-2029 AI can fully automate coding (do any coding task the best AGI-company engineer does, faster and cheaper). - Resolve by: 2029-12-31 - Verdict: Open - URL: https://sipoly.si/claim/kokotajlo-full-coding-automation-2029-median/ ### Most white-collar tasks automated within 12 to 18 months - Claimant: Mustafa Suleyman - Said: Feb 12, 2026 - Quote: "said in an interview with the Financial Times that he predicts most, if not every, task in white-collar fields will be automated by AI within the next year or year and a half." - Source: https://www.businessinsider.com/microsoft-ai-ceo-mustafa-suleyman-white-collar-tasks-automation-prediction-2026-2 - Resolution criteria (our interpretation): Our interpretation: by 12 Aug 2027, AI can fully perform most tasks of computer-based professional jobs (lawyers, accountants, project managers, marketers), as shown by broad deployment or independent evaluations. - Resolve by: 2027-08-12 - Verdict: Open - URL: https://sipoly.si/claim/suleyman-white-collar-automated-18m/ ### First clinical trials of its AI-designed drugs by end of 2026 - Claimant: Isomorphic Labs - Said: Jan 20, 2026 - Quote: "Isomorphic Labs, which uses artificial intelligence for drug discovery, expects to have its first clinical trials by the end of 2026, founder and CEO Demis Hassabis said on Tuesday." - Source: https://www.channelnewsasia.com/business/google-backed-isomorphic-labs-delays-clinical-trial-timeline-5871841 - Resolution criteria (our interpretation): Our interpretation: Isomorphic Labs starts a human clinical trial of one of its drug candidates by 31 Dec 2026. - Resolve by: 2026-12-31 - Verdict: Pending - URL: https://sipoly.si/claim/isomorphic-trials-end-2026/ ### Median 2030 for full coding automation (AI Futures Model) - Claimant: Eli Lifland - Said: Dec 31, 2025 - Quote: "With my parameters, it predicts a median of 2030 for full coding automation" - Source: https://x.com/eli_lifland/status/2006186170577817612 - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2030, AI can fully automate the coding work of top AI-company engineers (the AI 2027 "superhuman coder" milestone). - Resolve by: 2030-12-31 - Verdict: Open - URL: https://sipoly.si/claim/lifland-coding-automation-2030/ ### "Probably" by Dec 2026 the recursive self-improvement loop on algorithms will be closed - Claimant: David "davidad" Dalrymple - Said: Dec 20, 2025 - Quote: "I would guess that by December 2026 the RSI loop on algorithms will probably be closed, resulting in another inflection point to an even faster pace, perhaps around 70-80 day doubling time." - Source: https://x.com/davidad/status/2002412311999598604 - Resolution criteria (our interpretation): Our interpretation: correct if by 31 Dec 2026 (1) METR or an equivalent tracker reports that the doubling time of frontier 50% time horizons over 2026 is about 80 days or less, and (2) at least one frontier lab states publicly that AI systems now do most of its algorithmic research work. Wrong if neither holds; partly if only one does. - Resolve by: 2026-12-31 - Verdict: Open - URL: https://sipoly.si/claim/davidad-rsi-loop-dec-2026/ ### 2025 will be seen as the year of "peak bubble" for generative AI - Claimant: Gary Marcus - Said: Dec 20, 2025 - Quote: "2025 will be known as the year of the peak bubble , and also the moment at which Wall Street began to lose confidence in generative AI." - Source: https://garymarcus.substack.com/p/six-or-seven-predictions-for-ai-2026 - Resolution criteria (our interpretation): Our interpretation: by end-2026, major AI-exposed equity indices or AI company valuations are materially below their 2025 peaks. - Resolve by: 2026-12-31 - Verdict: Open - URL: https://sipoly.si/claim/marcus-peak-bubble/ ### No country takes a decisive lead in the GenAI race in 2026 - Claimant: Gary Marcus - Said: Dec 20, 2025 - Quote: "No country will take a decisive lead in the GenAI “race” ." - Source: https://garymarcus.substack.com/p/six-or-seven-predictions-for-ai-2026 - Resolution criteria (our interpretation): "Decisive lead" undefined; our interpretation: frontier-model leaderboards at end-2026 still show top models from more than one country within a small margin. - Resolve by: 2026-12-31 - Verdict: Open - URL: https://sipoly.si/claim/marcus-no-decisive-lead/ ### AI will be an election issue in the 2026 US midterms - Claimant: Gary Marcus - Said: Dec 20, 2025 - Quote: "In the midterms, AI will be an election issue for first time." - Source: https://garymarcus.substack.com/p/six-or-seven-predictions-for-ai-2026 - Resolution criteria (our interpretation): Our interpretation: AI appears among top issues in national exit polls or major-party campaign platforms for Nov 2026. - Resolve by: 2026-11-30 - Verdict: Open - URL: https://sipoly.si/claim/marcus-ai-midterms-issue/ ### Home humanoid robots "all demo and very little product" in 2026 - Claimant: Gary Marcus - Said: Dec 20, 2025 - Quote: "Human domestic robots like Optimus and Figure will be all demo and very little product." - Source: https://garymarcus.substack.com/p/six-or-seven-predictions-for-ai-2026 - Resolution criteria (our interpretation): Our interpretation: fewer than ~10,000 general-purpose humanoid robots delivered to private homes worldwide in 2026. - Resolve by: 2026-12-31 - Verdict: Open - URL: https://sipoly.si/claim/marcus-domestic-robots-2026/ ### No AGI in 2026 (or 2027) - Claimant: Gary Marcus - Said: Dec 20, 2025 - Quote: "We won’t get to AGI in 2026 (or 7)." - Source: https://garymarcus.substack.com/p/six-or-seven-predictions-for-ai-2026 - Resolution criteria (our interpretation): No system broadly accepted as AGI by 31 Dec 2027. - Resolve by: 2027-12-31 - Verdict: Open - URL: https://sipoly.si/claim/marcus-no-agi-2026-27/ ### A system that learns as well as a human, then becomes superhuman: "5 to 20" years - Claimant: Ilya Sutskever - Said: Nov 25, 2025 - Quote: "I think like 5 to 20." - Source: https://www.dwarkesh.com/p/ilya-sutskever-2 - Resolution criteria (our interpretation): Our interpretation: correct if such a system arrives between Nov 2030 and Nov 2045; wrong if before Nov 2030 or not by Nov 2045. - Resolve by: 2045-11-25 - Verdict: Open - URL: https://sipoly.si/claim/sutskever-superhuman-learner-5-20y/ ### "Something pretty clearly superhuman in most respects by end of 2027" - Claimant: Miles Brundage - Said: Nov 23, 2025 - Quote: "very roughly, something pretty clearly superhuman in most respects by end of 2027 + also very big stuff before then" - Source: https://x.com/Miles_Brundage/status/1992457550135275548 - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2027, an AI system outperforms top human experts in most cognitive domains, per broad expert consensus. - Resolve by: 2027-12-31 - Verdict: Open - URL: https://sipoly.si/claim/brundage-superhuman-by-2027/ ### Strong AGI expected "in the next 5-10 years" (from Nov 2025) - Claimant: Daniel Kokotajlo - Said: Nov 12, 2025 - Quote: "The companies seem to think strong AGI is just a few years away, and while I'm not as bullish as they are, I do expect it to happen in the next 5-10 years." - Source: https://x.com/DKokotajlo/status/1988697269584163255 - Resolution criteria (our interpretation): Our interpretation: by 12 Nov 2035, an AI system broadly accepted as strong AGI (able to do essentially all cognitive work humans do) exists. - Resolve by: 2035-11-12 - Verdict: Open - URL: https://sipoly.si/claim/kokotajlo-strong-agi-5-10y/ ### 1.6% of recently approved drug sales in 2027 will come from AI-discovered drugs (5% by 2030) - Claimant: LEAP expert panel - Said: Nov 10, 2025 - Quote: "Experts predict that 1.6% of recently approved drug sales in 2027 will come from AI-discovered drugs" - Source: https://leap.forecastingresearch.org/reports/wave2 - Resolution criteria (our interpretation): Our interpretation: correct if credible 2027 analyses put the share of sales of recently approved drugs attributable to AI-discovered drugs at 0.8-3%. - Resolve by: 2028-12-31 - Verdict: Open - URL: https://sipoly.si/claim/leap-ai-drug-sales-1-6pct-2027/ ### ~30% of physics, materials science and medicine papers will engage AI by 2030 (from 3% in 2022) - Claimant: LEAP expert panel - Said: Nov 10, 2025 - Quote: "Experts predict a 10x increase (from 3% to ~30%) in AI-engaged papers in Physics, 10 Materials Science, 11 and Medicine, 12 between 2022 and 2030." - Source: https://leap.forecastingresearch.org/reports/wave2 - Resolution criteria (our interpretation): Our interpretation: correct if the Duede et al.-style measure of AI-engaged papers reaches roughly 20-30% in these fields in 2030. - Resolve by: 2031-12-31 - Verdict: Open - URL: https://sipoly.si/claim/leap-ai-engaged-papers-30pct-2030/ ### AI will use 7% of US electricity in 2030 - Claimant: LEAP expert panel - Said: Nov 10, 2025 - Quote: "That rises to 7% of all electricity consumption in 2030" - Source: https://leap.forecastingresearch.org/reports/wave2 - Resolution criteria (our interpretation): Our interpretation: correct if credible 2030 estimates put AI training plus inference at 5.5-8.5% of US electricity consumption. - Resolve by: 2031-12-31 - Verdict: Open - URL: https://sipoly.si/claim/leap-ai-electricity-7pct-2030/ ### AI will use 4% of US electricity in 2027 (7% in 2030) - Claimant: LEAP expert panel - Said: Nov 10, 2025 - Quote: "The median expert predicts that 4% of U.S. electricity consumption will be used for training and deploying AI systems in 2027." - Source: https://leap.forecastingresearch.org/reports/wave2 - Resolution criteria (our interpretation): Our interpretation: correct if credible 2027 estimates (EIA, LBNL, IEA, Epoch) put AI training plus inference at 3-5% of US electricity consumption. - Resolve by: 2028-12-31 - Verdict: Open - URL: https://sipoly.si/claim/leap-ai-electricity-4pct-2027/ ### 20%: AI solves or substantially helps solve a Millennium Prize Problem by 2030 - Claimant: LEAP expert panel - Said: Nov 10, 2025 - Quote: "Experts estimate a 10% chance that AI will solve or substantially assist in solving a Millennium Prize Problem by 2027, 2 up to 20% by 2030" - Source: https://leap.forecastingresearch.org/reports/wave2 - Resolution criteria (our interpretation): Our interpretation: YES if by 31 Dec 2030 a Millennium Prize Problem is accepted as solved, with AI credited as solving it or substantially assisting. - Resolve by: 2030-12-31 - Verdict: Open - URL: https://sipoly.si/claim/leap-millennium-prize-2030-20pct/ ### 10%: AI solves or substantially helps solve a Millennium Prize Problem by 2027 - Claimant: LEAP expert panel - Said: Nov 10, 2025 - Quote: "Experts estimate a 10% chance that AI will solve or substantially assist in solving a Millennium Prize Problem by 2027" - Source: https://leap.forecastingresearch.org/reports/wave2 - Resolution criteria (our interpretation): Our interpretation: YES if by 31 Dec 2027 a Millennium Prize Problem is accepted as solved, with AI credited as solving it or substantially assisting. - Resolve by: 2027-12-31 - Verdict: Open - URL: https://sipoly.si/claim/leap-millennium-prize-2027-10pct/ ### "Pretty confident" of AI that makes more significant discoveries in 2028 and beyond - Claimant: OpenAI - Said: Nov 6, 2025 - Quote: "In 2028 and beyond, we are pretty confident we will have systems that can make more significant discoveries" - Source: https://openai.com/index/ai-progress-and-recommendations/ - Resolution criteria (our interpretation): Our interpretation: by the end of 2028, an AI system is credited with a significant scientific discovery (publishable in a top venue on its own merits, with AI as the main contributor). - Resolve by: 2028-12-31 - Verdict: Open - URL: https://sipoly.si/claim/openai-significant-discoveries-2028/ ### In 2026 AI will be able to make "very small discoveries" - Claimant: OpenAI - Said: Nov 6, 2025 - Quote: "In 2026, we expect AI to be capable of making very small discoveries." - Source: https://openai.com/index/ai-progress-and-recommendations/ - Resolution criteria (our interpretation): Our interpretation: during 2026, AI systems are credited (by the discoverers, in public write-ups) with at least one small but genuinely new scientific or mathematical result. - Resolve by: 2026-12-31 - Verdict: Open - URL: https://sipoly.si/claim/openai-small-discoveries-2026/ ### An automated "intern-level research assistant" by September 2026 - Claimant: OpenAI - Said: Oct 28, 2025 - Quote: "we think it is plausible that by September of next year, we have an intern-level AI research assistant" - Source: https://www.engadget.com/2251859/openai-says-it-reached-its-goal-of-creating-an-automated-research-intern/ - Resolution criteria (our interpretation): Our interpretation: by 30 Sep 2026, OpenAI has, by its own account, a system it calls an automated research intern doing well-defined research tasks under supervision. - Resolve by: 2026-09-30 - Verdict: Kept - URL: https://sipoly.si/claim/openai-research-intern-sep-2026/ ### A fully automated "legitimate AI researcher" by 2028 - Claimant: OpenAI - Said: Oct 28, 2025 - Quote: "internally, OpenAI is tracking toward achieving an intern-level research assistant by September 2026 and a fully automated “legitimate AI researcher” by 2028" - Source: https://techcrunch.com/2025/10/28/sam-altman-says-openai-will-have-a-legitimate-ai-researcher-by-2028/ - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2028, OpenAI has a system that autonomously carries out larger research projects, publicly demonstrated or credibly described. - Resolve by: 2028-12-31 - Verdict: Pending - URL: https://sipoly.si/claim/openai-legit-ai-researcher-2028/ ### "Only two years away from the first AI feature" (Oct 2025) - Claimant: Paul Schrader - Said: Oct 24, 2025 - Quote: "I think we’re only two years away from the first AI feature." - Source: https://www.vanityfair.com/hollywood/story/paul-schrader-ai-movie-interview - Resolution criteria (our interpretation): Our interpretation: correct if by 24 Oct 2027 a feature-length film (at least 75 minutes) made predominantly with generative AI gets a commercial theatrical or major-streamer release. - Resolve by: 2027-10-24 - Verdict: Open - URL: https://sipoly.si/claim/schrader-first-ai-feature-2y/ ### Tesla robotaxi in "eight to ten metro areas" by end of 2025 - Claimant: Tesla - Said: Oct 22, 2025 - Quote: "We do expect to be operating robotaxi in, I think, about eight to ten metro areas by the end of the year. It depends on various regulatory approvals." - Source: https://www.fool.com/earnings/call-transcripts/2025/10/22/tesla-tsla-q3-2025-earnings-call-transcript/ - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2025, Tesla operates a public robotaxi service in at least 8 US metro areas. - Resolve by: 2025-12-31 - Verdict: Broken - URL: https://sipoly.si/claim/tesla-robotaxi-8-10-metros-2025/ ### 10% (and rising) chance that Grok 5 achieves AGI - Claimant: Elon Musk - Said: Oct 18, 2025 - Quote: "My estimate of the probability of Grok 5 achieving AGI is now at 10% and rising" - Source: https://x.com/elonmusk/status/1979431839824777673 - Resolution criteria (our interpretation): Our interpretation: YES if Grok 5, once released, is broadly recognized as AGI by independent experts (not just by xAI). Scored at the end of 2026, or once Grok 5 has been out for 6 months if later. - Resolve by: 2026-12-31 - Verdict: Open - URL: https://sipoly.si/claim/musk-grok5-agi-10pct/ ### Truly capable AI agents are about "a decade" away, not "the year of agents" - Claimant: Andrej Karpathy - Said: Oct 17, 2025 - Quote: "I feel like the problems are tractable, they’re surmountable, but they’re still difficult. If I just average it out, it just feels like a decade to me." - Source: https://www.dwarkesh.com/p/andrej-karpathy - Resolution criteria (our interpretation): Our interpretation: "about a decade" means agents that can be hired like a reliable employee or intern for general knowledge work do not arrive before about Oct 2030 and do arrive by about Oct 2035. Scored at the end of the window. - Resolve by: 2035-10-17 - Verdict: Open - URL: https://sipoly.si/claim/karpathy-agents-decade/ ### Fully driverless Waymo rides in London in 2026 - Claimant: Waymo - Said: Oct 2025 - Quote: "We’re bringing our fully autonomous ride-hailing service across the pond, where we intend to offer rides – with no human behind the wheel – in 2026." - Source: https://waymo.com/blog/2025/10/hello-london-your-waymo-ride-is-arriving - Resolution criteria (our interpretation): Our interpretation: members of the public can take rides with no human behind the wheel in London by 31 Dec 2026. - Resolve by: 2026-12-31 - Verdict: Pending - URL: https://sipoly.si/claim/waymo-london-2026/ ### Models will work autonomously for full 8-hour days by mid-2026 - Claimant: Julian Schrittwieser - Said: Sep 27, 2025 - Quote: "Models will be able to autonomously work for full days (8 working hours) by mid-2026." - Source: https://www.julian.ac/blog/2025/09/27/failing-to-understand-the-exponential-again/ - Resolution criteria (our interpretation): Our interpretation: by 30 Jun 2026, METR (or an equivalent evaluation) reports a frontier model with a 50% time horizon of at least 8 hours. - Resolve by: 2026-06-30 - Verdict: Correct - URL: https://sipoly.si/claim/schrittwieser-full-day-autonomy-mid-2026/ ### By end-2027, models will frequently outperform experts on many tasks - Claimant: Julian Schrittwieser - Said: Sep 27, 2025 - Quote: "By the end of 2027, models will frequently outperform experts on many tasks." - Source: https://www.julian.ac/blog/2025/09/27/failing-to-understand-the-exponential-again/ - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2027, a model clearly beats (not just ties) industry experts on a majority of tasks in GDPval or an equivalent evaluation. - Resolve by: 2027-12-31 - Verdict: Open - URL: https://sipoly.si/claim/schrittwieser-outperform-experts-2027/ ### At least one model will match human experts across many industries before end-2026 - Claimant: Julian Schrittwieser - Said: Sep 27, 2025 - Quote: "At least one model will match the performance of human experts across many industries before the end of 2026." - Source: https://www.julian.ac/blog/2025/09/27/failing-to-understand-the-exponential-again/ - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2026, a model wins or ties at least 50% of comparisons against industry experts on GDPval (or an equivalent cross-industry evaluation). - Resolve by: 2026-12-31 - Verdict: Open - URL: https://sipoly.si/claim/schrittwieser-expert-parity-2026/ ### "Very likely" fully operational 1 GW data centers before mid-2026 - Claimant: Peter Wildeford - Said: Sep 25, 2025 - Quote: "However, it is very likely we will see fully operational 1GW data centers before mid-2026 ." - Source: https://peterwildeford.substack.com/p/openai-nvidia-and-oracle-breaking - Resolution criteria (our interpretation): Our interpretation: by 30 Jun 2026, at least one AI data center reaches 1 GW of operational power, per Epoch AI or similar trackers. - Resolve by: 2026-06-30 - Verdict: Correct - URL: https://sipoly.si/claim/wildeford-1gw-datacenter-mid-2026/ ### "Real AI" (beyond today's chatbots) arrives in about 5 years - Claimant: David Krueger - Said: Sep 18, 2025 - Quote: "No. We’re not there yet. But the real AI is coming. I think we’ve got about 5 years." - Source: https://therealartificialintelligence.substack.com/p/announcing-the-real-ai-a-blog - Resolution criteria (our interpretation): Our interpretation: AI that can broadly replace human cognitive work arrives by about Sep 2030 (we allow a year either way). - Resolve by: 2031-09-18 - Verdict: Open - URL: https://sipoly.si/claim/krueger-real-ai-5-years/ ### By 2030 AI will autonomously fix issues, implement features and solve hard scientific programming problems - Claimant: Epoch AI - Said: Sep 2025 - Quote: "By 2030, on current trends, AI will be able to autonomously fix issues, implement features, and solve difficult (but well-defined) scientific programming problems." - Source: https://epoch.ai/publications/what-will-ai-look-like-in-2030 - Resolution criteria (our interpretation): Our interpretation: by end-2030, AI agents routinely resolve real-world software issues and implement features end to end without human code edits (e.g. SWE-bench-style benchmarks saturated, plus documented production use), and solve well-specified scientific programming tasks. - Resolve by: 2030-12-31 - Verdict: Open - URL: https://sipoly.si/claim/epoch-autonomous-swe-2030/ ### Frontier training clusters will cost over $100B by 2030, enabling ~1e29 FLOP runs - Claimant: Epoch AI - Said: Sep 2025 - Quote: "On current trends, the clusters used for training frontier AI would cost over $100B by 2030. Such clusters could support training runs of about 10^29 FLOP" - Source: https://epoch.ai/publications/what-will-ai-look-like-in-2030 - Resolution criteria (our interpretation): Our interpretation: correct if by end-2030 at least one frontier training cluster has a reported cost above $100B, and the largest known training run is within 10x of 1e29 FLOP. - Resolve by: 2030-12-31 - Verdict: Open - URL: https://sipoly.si/claim/epoch-100b-clusters-1e29-2030/ ### Few if any drugs approved by 2030 will have benefited from today's AI tools - Claimant: Epoch AI - Said: Sep 2025 - Quote: "we expect that few if any of the drugs approved for sale by 2030 will have benefited from today’s AI tools, let alone those of 2030." - Source: https://epoch.ai/publications/what-will-ai-look-like-in-2030 - Resolution criteria (our interpretation): Our interpretation: correct if, by end-2030, no more than a handful (≤3) of FDA/EMA-approved drugs are credibly described as discovered or designed with 2025-era AI tools. - Resolve by: 2030-12-31 - Verdict: Open - URL: https://sipoly.si/claim/epoch-few-ai-drugs-by-2030/ ### By 2030 many scientific fields will have AI assistants comparable to today's coding assistants - Claimant: Epoch AI - Said: Sep 2025 - Quote: "By 2030, we predict that many scientific domains will have AI assistants comparable to coding assistants for software engineers today." - Source: https://epoch.ai/publications/what-will-ai-look-like-in-2030 - Resolution criteria (our interpretation): Our interpretation: by 2030, AI research assistants are in routine use (comparable to 2025 coding assistants) in most major scientific fields. - Resolve by: 2030-12-31 - Verdict: Open - URL: https://sipoly.si/claim/epoch-science-assistants-2030/ ### "In five years" unemployment levels never seen before: "Not talking about 10% ... but 99%" - Claimant: Roman Yampolskiy - Said: Sep 4, 2025 - Quote: "In five years, we’re looking at levels of unemployment we’ve never seen before" - Source: https://www.businessinsider.com/ai-safety-pioneer-predicts-ai-could-cause-99-unemployment-by-2030-2025-9 - Resolution criteria (our interpretation): Our interpretation: correct if by Sep 2030 the US unemployment rate exceeds its post-1948 record (14.8%). Full credit for his headline "99%" would need near-total unemployment. - Resolve by: 2030-09-04 - Verdict: Open - URL: https://sipoly.si/claim/yampolskiy-99pct-unemployment-5y/ ### AGI "likely to arrive by 2027" - Claimant: Roman Yampolskiy - Said: Sep 4, 2025 - Quote: "said artificial general intelligence — systems as capable as humans across domains — is likely to arrive by 2027." - Source: https://www.businessinsider.com/ai-safety-pioneer-predicts-ai-could-cause-99-unemployment-by-2030-2025-9 - Resolution criteria (our interpretation): Our interpretation: AI as capable as humans across domains exists by 31 Dec 2027, by broad expert recognition. - Resolve by: 2027-12-31 - Verdict: Open - URL: https://sipoly.si/claim/yampolskiy-agi-2027/ ### ~25%: 80%-reliability one-month task horizons by the start of 2028 - Claimant: Ryan Greenblatt - Said: Aug 2025 - Quote: "We're pretty unlikely (25%?) to see 80% reliability at 1 month tasks by the start of 2028." - Source: https://www.lesswrong.com/posts/2ssPfDpdrjaM2rMbn/my-agi-timeline-updates-from-gpt-5-and-2025-so-far-1 - Resolution criteria (our interpretation): Our interpretation: YES if by 1 Jan 2028 METR (or an equivalent evaluation) reports an 80%-reliability time horizon of at least one work-month (~167 hours). - Resolve by: 2028-01-01 - Verdict: Open - URL: https://sipoly.si/claim/greenblatt-1-month-80pct-2028-25pct/ ### METR 50% time horizons of about two weeks around the start of 2028 - Claimant: Ryan Greenblatt - Said: Aug 2025 - Quote: "I expect doubling times of around 170 days on METR's task suite (or similar tasks) [3] over the next 2 years or so which implies we'll be hitting 2 week 50% reliability horizon lengths around the start of 2028." - Source: https://www.lesswrong.com/posts/2ssPfDpdrjaM2rMbn/my-agi-timeline-updates-from-gpt-5-and-2025-so-far-1 - Resolution criteria (our interpretation): Our interpretation: correct if METR's best published 50% time horizon first reaches about 2 work-weeks (~80 hours) between Jul 2027 and Jun 2028; wrong if it happens before Jul 2027 or is not reached by Jun 2028. - Resolve by: 2028-06-30 - Verdict: Open - URL: https://sipoly.si/claim/greenblatt-2-week-horizon-2028/ ### 45%: AI capable of fully automating AI R&D by the start of 2033 - Claimant: Ryan Greenblatt - Said: Aug 2025 - Quote: "I now would put ~15% probability by the start of 2029 and a 45% chance by the start of 2033." - Source: https://www.lesswrong.com/posts/2ssPfDpdrjaM2rMbn/my-agi-timeline-updates-from-gpt-5-and-2025-so-far-1 - Resolution criteria (our interpretation): Our interpretation: YES if by 1 Jan 2033 an AI system exists that could fully automate AI R&D. - Resolve by: 2033-01-01 - Verdict: Open - URL: https://sipoly.si/claim/greenblatt-ai-rd-automation-2033-45pct/ ### ~15%: AI capable of fully automating AI R&D by the start of 2029 - Claimant: Ryan Greenblatt - Said: Aug 2025 - Quote: "I now would put ~15% probability by the start of 2029 and a 45% chance by the start of 2033." - Source: https://www.lesswrong.com/posts/2ssPfDpdrjaM2rMbn/my-agi-timeline-updates-from-gpt-5-and-2025-so-far-1 - Resolution criteria (our interpretation): Our interpretation: YES if by 1 Jan 2029 an AI system exists that could fully automate AI R&D (do all the work of a frontier lab's research staff). - Resolve by: 2029-01-01 - Verdict: Open - URL: https://sipoly.si/claim/greenblatt-ai-rd-automation-2029-15pct/ ### Dojo 2 "operating at scale sometime next year" (2026) - Claimant: Tesla - Said: Jul 23, 2025 - Quote: "Dojo two, we expect to have Dojo two operating at scale sometime next year. With scale being somewhere around a hundred k h one hundred equivalent." - Source: https://www.fool.com/earnings/call-transcripts/2025/07/23/tesla-tsla-q2-2025-earnings-call-transcript/ - Resolution criteria (our interpretation): Our interpretation: Tesla's Dojo 2 training system runs at roughly 100k-H100-equivalent scale during 2026. - Resolve by: 2026-12-31 - Verdict: Withdrawn - URL: https://sipoly.si/claim/tesla-dojo2-at-scale-2026/ ### Owners able to add their cars to the Tesla robotaxi fleet in 2026, "confidently next year" - Claimant: Tesla - Said: Jul 23, 2025 - Quote: "But I guess next year is I'd say confidently next year. I'm not sure when next year, but confidently next year, people would be able to add or subtract their car to the Tesla, Inc. fleet." - Source: https://www.fool.com/earnings/call-transcripts/2025/07/23/tesla-tsla-q2-2025-earnings-call-transcript/ - Resolution criteria (our interpretation): Our interpretation: during 2026, Tesla lets private owners put their own cars into its paid robotaxi network. - Resolve by: 2026-12-31 - Verdict: Pending - URL: https://sipoly.si/claim/tesla-add-own-car-to-fleet-2026/ ### Optimus 3 production to start "at the beginning of next year" (2026) - Claimant: Tesla - Said: Jul 23, 2025 - Quote: "We will have prototypes of that in about three months, and it's subtooling in production. We will start production on that at the beginning of next year." - Source: https://www.fool.com/earnings/call-transcripts/2025/07/23/tesla-tsla-q2-2025-earnings-call-transcript/ - Resolution criteria (our interpretation): Our interpretation: Tesla starts production of Optimus 3 by 31 Mar 2026. - Resolve by: 2026-03-31 - Verdict: Broken - URL: https://sipoly.si/claim/tesla-optimus3-production-early-2026/ ### Unsupervised FSD for personally owned cars by end of 2025 "in certain geographies" - Claimant: Tesla - Said: Jul 23, 2025 - Quote: "I think it will be available for unsupervised personal use by the end of this year in certain geographies." - Source: https://www.fool.com/earnings/call-transcripts/2025/07/23/tesla-tsla-q2-2025-earnings-call-transcript/ - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2025, owners in at least one US region can use FSD without supervision in their own cars. - Resolve by: 2025-12-31 - Verdict: Broken - URL: https://sipoly.si/claim/tesla-unsupervised-personal-fsd-2025/ ### Tesla autonomous ride-hailing for "half the population of the US" by end of 2025 - Claimant: Tesla - Said: Jul 23, 2025 - Quote: "I think we will probably have autonomous ride-hailing in probably half the population of the US by the end of the year. That's at least our goal, subject to regulatory approvals." - Source: https://www.fool.com/earnings/call-transcripts/2025/07/23/tesla-tsla-q2-2025-earnings-call-transcript/ - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2025, Tesla autonomous ride-hailing is available in metro areas holding about half the US population (~165 million people). - Resolve by: 2025-12-31 - Verdict: Broken - URL: https://sipoly.si/claim/tesla-robotaxi-half-us-2025/ ### AI "world-class mathematicians" within a year, world-class programmers within one or two (Jul 2025) - Claimant: Eric Schmidt - Said: Jul 17, 2025 - Quote: "So, it's likely in my opinion that you're going to see world-class mathematicians emerge in the next one year that are AI based and world-class programmers that going to appear within the next one or two years." - Source: https://www.youtube.com/watch?v=qaPHK1fJL5s - Resolution criteria (our interpretation): Our interpretation: correct if by 17 Jul 2026 an AI system is recognized by leading mathematicians as doing world-class research mathematics (e.g. original results accepted in top journals) AND by 17 Jul 2027 AI systems perform at world-class level in programming competitions and real-world engineering; partly if only one half holds. - Resolve by: 2027-07-17 - Verdict: Open - URL: https://sipoly.si/claim/schmidt-ai-world-class-mathematicians-1y/ ### Prometheus, a 1 GW AI supercluster, online in 2026 - Claimant: Meta - Said: Jul 14, 2025 - Quote: "We're actually building several multi-GW clusters. We're calling the first one Prometheus and it's coming online in '26." - Source: https://www.datacenterdynamics.com/en/news/meta-to-invest-hundreds-of-billions-of-dollars-into-compute-to-build-superintelligence-with-several-multi-gw-data-center-clusters/ - Resolution criteria (our interpretation): Our interpretation: correct if during 2026 Meta's Prometheus cluster (New Albany, Ohio) is operating at about 1 GW or more, per Meta or Epoch AI; partly if it comes online in 2026 but well below 1 GW; wrong if it is not online in 2026. - Resolve by: 2026-12-31 - Verdict: Pending - URL: https://sipoly.si/claim/meta-prometheus-1gw-2026/ ### Hyperion data centre: 2 GW online by 2030, scaling to 5 GW - Claimant: Meta - Said: Jul 14, 2025 - Quote: "Gabriel says Meta plans to bring two gigawatts of data center capacity online by 2030 with Hyperion, but that it would scale to five gigawatts in several years." - Source: https://techcrunch.com/2025/07/14/mark-zuckerberg-says-meta-is-building-a-5gw-ai-data-center/ - Resolution criteria (our interpretation): Our interpretation: Meta's Hyperion campus has at least 2 GW of data-centre capacity operational by 31 Dec 2030 (per Meta, utility filings or Epoch AI satellite tracking). - Resolve by: 2030-12-31 - Verdict: Pending - URL: https://sipoly.si/claim/meta-hyperion-2gw-2030/ ### Within five years AI will be able to do 80% of 80% of all jobs - Claimant: Vinod Khosla - Said: Jun 2025 - Quote: "“Within the next five years, any economically valuable job humans can do, AI will be able to do 80% of it…80% of all jobs can be done by an AI.”" - Source: https://fortune.com/2025/07/01/silicon-valley-investor-vinod-khosla-ai-job-prediction-interview/ - Resolution criteria (our interpretation): Our interpretation: by mid-2030, AI systems can perform about 80% of the tasks in about 80% of economically valuable jobs (capability, not necessarily deployment). - Resolve by: 2030-07-01 - Verdict: Open - URL: https://sipoly.si/claim/khosla-ai-80pct-of-jobs-5y/ ### 33% of enterprise software apps will include agentic AI by 2028; 15% of daily work decisions autonomous - Claimant: Gartner - Said: Jun 25, 2025 - Quote: "Gartner predicts at least 15% of day-to-day work decisions will be made autonomously through agentic AI by 2028, up from 0% in 2024. In addition, 33% of enterprise software applications will include agentic AI by 2028, up from less than 1% in 2024." - Source: https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027 - Resolution criteria (our interpretation): Our interpretation: correct if Gartner's own (or comparable analyst) 2028 data show at least 30% of enterprise software apps include agentic AI and at least 15% of day-to-day work decisions are made autonomously. - Resolve by: 2028-12-31 - Verdict: Open - URL: https://sipoly.si/claim/gartner-agentic-33pct-apps-2028/ ### Over 40% of agentic AI projects canceled by end of 2027 - Claimant: Gartner - Said: Jun 25, 2025 - Quote: "Over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls" - Source: https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027 - Resolution criteria (our interpretation): Our interpretation: Gartner or other credible surveys covering 2027 report that more than 40% of enterprise agentic-AI projects were canceled. - Resolve by: 2027-12-31 - Verdict: Open - URL: https://sipoly.si/claim/gartner-agentic-40pct-canceled-2027/ ### Zoox to start public robotaxi rides in Las Vegas later in 2025 - Claimant: Zoox - Said: Jun 18, 2025 - Quote: "Zoox is gearing up to start public rides in Las Vegas later this year, with San Francisco to follow." - Source: https://www.cnbc.com/2025/06/18/amazon-zoox-robotaxi.html - Resolution criteria (our interpretation): Our interpretation: members of the public can hail Zoox robotaxi rides in Las Vegas by 31 Dec 2025. - Resolve by: 2025-12-31 - Verdict: Kept - URL: https://sipoly.si/claim/zoox-vegas-public-rides-2025/ ### Generative AI will reduce Amazon's total corporate workforce "in the next few years" - Claimant: Amazon - Said: Jun 2025 - Quote: "It’s hard to know exactly where this nets out over time, but in the next few years, we expect that this will reduce our total corporate workforce as we get efficiency gains from using AI extensively across the company." - Source: https://www.aboutamazon.com/news/company-news/amazon-ceo-andy-jassy-on-generative-ai - Resolution criteria (our interpretation): Our interpretation: Amazon's corporate (non-warehouse) headcount falls within about three years (by mid-2028), with AI cited as a driver. - Resolve by: 2028-06-30 - Verdict: Open - URL: https://sipoly.si/claim/amazon-ai-shrinks-corporate-workforce/ ### 2027 may see robots that can do tasks in the real world - Claimant: Sam Altman - Said: Jun 10, 2025 - Quote: "2027 may see the arrival of robots that can do tasks in the real world." - Source: https://blog.samaltman.com/the-gentle-singularity - Resolution criteria (our interpretation): Our interpretation: by end of 2027, general-purpose robots are commercially deployed doing varied real-world tasks outside controlled demos. - Resolve by: 2027-12-31 - Verdict: Open - URL: https://sipoly.si/claim/altman-2027-robots/ ### 2026 will likely see AI systems that figure out novel insights - Claimant: Sam Altman - Said: Jun 10, 2025 - Quote: "2026 will likely see the arrival of systems that can figure out novel insights." - Source: https://blog.samaltman.com/the-gentle-singularity - Resolution criteria (our interpretation): Our interpretation: during 2026 an AI system is credited, in a peer-reviewed paper or by recognised domain experts, as the principal originator of a genuinely new scientific or mathematical result. - Resolve by: 2026-12-31 - Verdict: Open - URL: https://sipoly.si/claim/altman-2026-novel-insights/ ### AGI "slightly after" 2030, not quite by 2030 - Claimant: Sundar Pichai - Said: Jun 5, 2025 - Quote: "I don’t think we’ll quite get there by 2030, so my sense is it’s slightly after that" - Source: https://lexfridman.com/sundar-pichai-transcript/ - Resolution criteria (our interpretation): Our interpretation: correct if AGI (as broadly recognized) has not arrived by the end of 2030 but does arrive by about 2033. - Resolve by: 2033-12-31 - Verdict: Open - URL: https://sipoly.si/claim/pichai-agi-slightly-after-2030/ ### Maybe 50% chance of AGI in five to ten years - Claimant: Demis Hassabis - Said: Jun 4, 2025 - Quote: "In the next five to 10 years, there’s maybe a 50 percent chance that we'll have what we define as AGI." - Source: https://www.wired.com/story/google-deepminds-ceo-demis-hassabis-thinks-ai-will-make-humans-less-selfish/ - Resolution criteria (our interpretation): Our interpretation: by 4 Jun 2035, a system exhibits all the cognitive capabilities humans have, per broad expert consensus. - Resolve by: 2035-06-04 - Verdict: Open - URL: https://sipoly.si/claim/hassabis-agi-50pct-5-10y/ ### 50/50 that AI learns on the job like a human for any white-collar work by 2032 - Claimant: Dwarkesh Patel - Said: Jun 2, 2025 - Quote: "AI learns on the job as easily, organically, seamlessly, and quickly as a human, for any white collar work." - Source: https://www.dwarkesh.com/p/timelines-june-2025 - Resolution criteria (our interpretation): Our interpretation: YES if by 31 Dec 2032 deployed AI systems show human-like continual learning on the job across white-collar work. - Resolve by: 2032-12-31 - Verdict: Open - URL: https://sipoly.si/claim/dwarkesh-on-the-job-learning-2032/ ### 50/50 that AI does small-business taxes end-to-end by 2028 - Claimant: Dwarkesh Patel - Said: Jun 2, 2025 - Quote: "AI can do taxes end-to-end for my small business as well as a competent general manager could in a week: including chasing down all the receipts on different websites, finding all the missing pieces, emailing back and forth with anyone we need to hassle for invoices, filling out the form, and sending it to the IRS: 2028" - Source: https://www.dwarkesh.com/p/timelines-june-2025 - Resolution criteria (our interpretation): Our interpretation: YES if by 31 Dec 2028 a generally available AI agent can do this end-to-end task as described. - Resolve by: 2028-12-31 - Verdict: Open - URL: https://sipoly.si/claim/dwarkesh-ai-taxes-2028/ ### AI better, faster and cheaper at any screen-based job "by next year" - Claimant: Emad Mostaque - Said: Jun 2025 - Quote: "For any job that you can do on the other side of a screen, an AI will probably be able to do it better, faster, and cheaper by next year." - Source: https://www.youtube.com/watch?v=lLzYKqcKM34 - Resolution criteria (our interpretation): Our interpretation: correct if by 31 Dec 2026 AI can do essentially any remote (screen-based) job better, faster and more cheaply than the humans doing it. - Resolve by: 2026-12-31 - Verdict: Open - URL: https://sipoly.si/claim/mostaque-screen-jobs-by-2026/ ### Digital superintelligence "this year" (2025), or "next year for sure" - Claimant: Elon Musk - Said: Jun 2025 - Quote: "I think we're quite close to digital superintelligence. It may happen this year and if it doesn't happen this year, next year for sure. A digital superintelligence defined as smarter than any human at anything." - Source: https://www.youtube.com/watch?v=cFIlta1GkiE - Resolution criteria (our interpretation): Our interpretation: correct if by 31 Dec 2026 an AI system is broadly acknowledged as smarter than any human at anything (superhuman in essentially every cognitive domain). - Resolve by: 2026-12-31 - Verdict: Open - URL: https://sipoly.si/claim/musk-digital-superintelligence-2026/ ### AI could wipe out half of entry-level white-collar jobs and push unemployment to 10-20% within 1-5 years - Claimant: Dario Amodei - Said: May 28, 2025 - Quote: "AI could wipe out half of all entry-level white-collar jobs — and spike unemployment to 10-20% in the next one to five years, Amodei told us in an interview from his San Francisco office." - Source: https://www.axios.com/2025/05/28/ai-jobs-white-collar-unemployment-anthropic - Resolution criteria (our interpretation): Our interpretation: by 28 May 2030, US unemployment (BLS U-3) reaches 10% or more at some point, with AI widely credited, or entry-level white-collar employment falls by about half. Hedged ("could"). - Resolve by: 2030-05-28 - Verdict: Open - URL: https://sipoly.si/claim/amodei-white-collar-half-1-5y/ ### AI able to replace more than half of humans "within 5 years" (from May 2025) - Claimant: Scott Alexander - Said: May 23, 2025 - Quote: "I think AI will be able to replace >50% of humans within 5 years." - Source: https://x.com/slatestarcodex/status/1925942621068792291 - Resolution criteria (our interpretation): Our interpretation: by 23 May 2030, AI can do the jobs of more than half of the human workforce at lower cost (capability, not necessarily deployment). - Resolve by: 2030-05-23 - Verdict: Open - URL: https://sipoly.si/claim/scott-alexander-replace-half-5y/ ### "Almost guaranteed": a drop-in white-collar worker within five years, "very likely in two" (May 2025) - Claimant: Sholto Douglas - Said: May 22, 2025 - Quote: "But the one that I feel we're almost guaranteed to get—this is a strong statement to make—is one where at the very least, you get a drop-in white collar worker at some point in the next five years. I think it's very likely in two, but it seems almost overdetermined in five." - Source: https://www.dwarkesh.com/p/sholto-trenton-2 - Resolution criteria (our interpretation): Our interpretation: correct if by 22 May 2030 AI agents can be deployed as drop-in replacements for typical white-collar workers (doing essentially any white-collar job end to end), as broadly recognized by independent observers; a developer's own claim does not count on its own. - Resolve by: 2030-05-22 - Verdict: Open - URL: https://sipoly.si/claim/sholto-white-collar-automatable-2028/ ### First 200 MW of Stargate UAE expected to go live in 2026 - Claimant: OpenAI - Said: May 22, 2025 - Quote: "A 1GW Stargate UAE cluster in Abu Dhabi with 200MW expected to go live in 2026" - Source: https://openai.com/index/introducing-stargate-uae/ - Resolution criteria (our interpretation): Our interpretation: the first 200 MW of Stargate UAE is operational by 31 Dec 2026. - Resolve by: 2026-12-31 - Verdict: Pending - URL: https://sipoly.si/claim/openai-stargate-uae-200mw-2026/ ### No halving of US labor-force participation and no 15% GDP growth year by 2029 - Claimant: Matthew Barnett - Said: May 18, 2025 - Quote: "I don't expect the US labor force participation rate to fall by more than 50% from its current level by 2029, or for US GDP growth to surpass 15% for any year in 2025-2029." - Source: https://x.com/MatthewJBar/status/1924205015490830473 - Resolution criteria (our interpretation): Our interpretation: correct if through 2029 the US labor force participation rate never falls below half its May 2025 level (~62.4%) and no calendar year 2025-2029 has real US GDP growth above 15%. - Resolve by: 2029-12-31 - Verdict: Open - URL: https://sipoly.si/claim/barnett-no-lfpr-collapse-2029/ ### Most code for Meta's AI efforts written by AI within 12 to 18 months - Claimant: Mark Zuckerberg - Said: Apr 29, 2025 - Quote: "I would guess that sometime in the next 12 to 18 months, we'll reach the point where most of the code that's going toward these efforts is written by AI." - Source: https://www.dwarkesh.com/p/mark-zuckerberg-2 - Resolution criteria (our interpretation): Our interpretation: by 29 Oct 2026, Meta states (or credible reporting shows) that most code for its AI-model efforts is AI-written, beyond autocomplete. - Resolve by: 2026-10-29 - Verdict: Open - URL: https://sipoly.si/claim/zuckerberg-most-code-12-18m/ ### Transformative economic and societal impacts of AI will take decades - Claimant: Arvind Narayanan - Said: Apr 15, 2025 - Quote: "we think that transformative economic and societal impacts will be slow (on the timescale of decades)" - Source: https://knightcolumbia.org/content/ai-as-normal-technology - Resolution criteria (our interpretation): Our interpretation: correct if by 15 Apr 2035 there has been no economy-wide AI transformation (e.g. no decade of US labor-productivity growth above 3%/yr, no AI-driven unemployment shock above 10%). - Resolve by: 2035-04-15 - Verdict: Open - URL: https://sipoly.si/claim/narayanan-kapoor-impacts-decades/ ### "In one year, the vast majority of programmers will be replaced by AI programmers" - Claimant: Eric Schmidt - Said: Apr 10, 2025 - Quote: "We believe, as an industry, that in one year, the vast majority of programmers will be replaced by AI programmers" - Source: https://san.com/cc/former-google-ceo-predicts-ai-will-replace-most-programmers-in-a-year/ - Resolution criteria (our interpretation): Our interpretation: correct if by Apr 2026 most professional programmer roles have been replaced by AI agents (e.g. US software-developer employment down by more than half). - Resolve by: 2026-04-10 - Verdict: Wrong - URL: https://sipoly.si/claim/schmidt-programmers-replaced-1y/ ### Data centre electricity demand will more than double to ~945 TWh by 2030 - Claimant: International Energy Agency - Said: Apr 10, 2025 - Quote: "It projects that electricity demand from data centres worldwide is set to more than double by 2030 to around 945 terawatt-hours (TWh), slightly more than the entire electricity consumption of Japan today." - Source: https://www.iea.org/news/ai-is-set-to-drive-surging-electricity-demand-from-data-centres-while-offering-the-potential-to-transform-how-the-energy-sector-works - Resolution criteria (our interpretation): Our interpretation: correct if IEA (or comparable) data put global data-centre electricity consumption in 2030 at 850-1,050 TWh. - Resolve by: 2031-12-31 - Verdict: Open - URL: https://sipoly.si/claim/iea-datacentre-945twh-2030/ ### Mid-2025: specialized coding and research agents begin to transform their professions - Claimant: AI Futures Project - Said: Apr 3, 2025 - Quote: "Meanwhile, out of public focus, more specialized coding and research agents are beginning to transform their professions." - Source: https://ai-2027.com/ - Resolution criteria (our interpretation): Our interpretation: correct if by late 2025 coding/research agents (e.g. Claude Code, Codex, Deep Research) are in widespread professional use and changing workflows. - Resolve by: 2025-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/ai2027-mid2025-coding-agents-transform/ ### Mid-2025: AI agents impressive in theory but unreliable in practice - Claimant: AI Futures Project - Said: Apr 3, 2025 - Quote: "The agents are impressive in theory (and in cherry-picked examples), but in practice unreliable." - Source: https://ai-2027.com/ - Resolution criteria (our interpretation): Our interpretation: correct if in mid/late 2025 general-purpose agents are widely reported as unreliable outside cherry-picked demos. - Resolve by: 2025-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/ai2027-mid2025-agents-unreliable/ ### A 10,000-person anti-AI protest in Washington, DC by late 2026 - Claimant: AI Futures Project - Said: Apr 3, 2025 - Quote: "Many people fear that the next wave of AIs will come for their jobs; there is a 10,000 person anti-AI protest in DC." - Source: https://ai-2027.com/ - Resolution criteria (our interpretation): Our interpretation: correct if a protest against AI with roughly 10,000 or more participants (credible media estimates) takes place in Washington, DC by 31 Dec 2026. - Resolve by: 2026-12-31 - Verdict: Open - URL: https://sipoly.si/claim/ai2027-10k-anti-ai-protest-dc-2026/ ### The stock market will be up 30% in 2026, led by AI companies - Claimant: AI Futures Project - Said: Apr 3, 2025 - Quote: "The stock market has gone up 30% in 2026, led by OpenBrain, Nvidia, and whichever companies have most successfully integrated AI assistants." - Source: https://ai-2027.com/ - Resolution criteria (our interpretation): Our interpretation: correct if the S&P 500 total return for calendar 2026 is at least +25%, with AI-linked companies leading. - Resolve by: 2026-12-31 - Verdict: Open - URL: https://sipoly.si/claim/ai2027-stock-market-up-30pct-2026/ ### Only about 30% chance of full remote-work automation within 10 years - Claimant: Ege Erdil - Said: Apr 2025 - Quote: "I still think full automation of remote work in 10 years is plausible, because it’s what we would predict if we straightforwardly extrapolate current rates of revenue growth and assume no slowdown. However, I would only give this outcome around 30% chance." - Source: https://epoch.ai/gradient-updates/the-case-for-multi-decade-ai-timelines - Resolution criteria (our interpretation): Our interpretation: YES if, by Apr 2035, AI can do essentially all remote (computer-based) jobs at or below human cost. - Resolve by: 2035-04-30 - Verdict: Open - URL: https://sipoly.si/claim/erdil-remote-work-automation-10y-30pct/ ### Open-weight reasoning model "in the coming months" - Claimant: OpenAI - Said: Mar 31, 2025 - Quote: "we are excited to release a powerful new open-weight language model with reasoning in the coming months" - Source: https://decrypt.co/312522/openai-release-open-weight-model-reasoning-capabilities - Resolution criteria (our interpretation): Open-weight reasoning model released by 31 Dec 2025. - Resolve by: 2025-12-31 - Verdict: Kept - URL: https://sipoly.si/claim/openai-open-weight-model/ ### Humanoid dexterity "will remain pathetic" compared to human hands through 2035 - Claimant: Rodney Brooks - Said: Mar 26, 2025 - Quote: "Humanoid Robots . Deployable dexterity will remain pathetic compared to human hands beyond 2036. Without new types of mechanical systems walking humanoids will remain too unsafe to be in close proximity to real humans." - Source: https://rodneybrooks.com/predictions-scorecard-2026-january-01/ - Resolution criteria (our interpretation): Our interpretation: correct if on 1 Jan 2036 no deployed humanoid robot has hand dexterity broadly comparable to human hands, and walking humanoids are not working in close proximity to people at scale. - Resolve by: 2036-01-01 - Verdict: Open - URL: https://sipoly.si/claim/brooks-2025-humanoid-dexterity/ ### Only Waymo and Zoox will matter for US self-driving through 2035 - Claimant: Rodney Brooks - Said: Mar 26, 2025 - Quote: "In the US the players that will determine whether self driving cars are successful or abandoned are #1 Waymo (Google) and #2 Zoox (Amazon). No one else matters." - Source: https://rodneybrooks.com/predictions-scorecard-2026-january-01/ - Resolution criteria (our interpretation): Our interpretation: correct if at 1 Jan 2036 Waymo and Zoox are the two largest US driverless ride services by rides or fleet, and no other company (e.g. Tesla) runs a comparable driverless service. - Resolve by: 2036-01-01 - Verdict: Open - URL: https://sipoly.si/claim/brooks-2025-waymo-zoox-only/ ### Waymo One ready for riders in Washington, D.C. in 2026 - Claimant: Waymo - Said: Mar 2025 - Quote: "It’s official: Waymo One, the world’s leading fully autonomous ride-hailing service, will be ready for riders in the nation’s capital on the Waymo One app in 2026." - Source: https://waymo.com/blog/2025/03/next-stop-for-waymo-one-washingtondc/ - Resolution criteria (our interpretation): Our interpretation: public riders can take fully autonomous Waymo rides in Washington, D.C. by 31 Dec 2026. - Resolve by: 2026-12-31 - Verdict: Pending - URL: https://sipoly.si/claim/waymo-dc-2026/ ### Within a decade, AI agents will independently do many software tasks that take humans days or weeks - Claimant: METR - Said: Mar 19, 2025 - Quote: "Extrapolating this trend predicts that, in under a decade, we will see AI agents that can independently complete a large fraction of software tasks that currently take humans days or weeks." - Source: https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ - Resolution criteria (our interpretation): Our interpretation: by 19 Mar 2035, METR (or equivalent) reports frontier agents with a 50% time horizon of at least one work-week (40 hours) on software tasks. - Resolve by: 2035-03-19 - Verdict: Open - URL: https://sipoly.si/claim/metr-week-long-tasks-decade/ ### Rubin Ultra (NVL576) in the second half of 2027 - Claimant: Nvidia - Said: Mar 18, 2025 - Quote: "Huang also announced Rubin Ultra, which will follow in the second half of 2027." - Source: https://arstechnica.com/ai/2025/03/nvidia-announces-rubin-ultra-and-feynman-ai-chips-for-2027-and-2028/ - Resolution criteria (our interpretation): Our interpretation: Nvidia ships Rubin Ultra (NVL576 systems) to customers by 31 Dec 2027. - Resolve by: 2027-12-31 - Verdict: Pending - URL: https://sipoly.si/claim/nvidia-rubin-ultra-h2-2027/ ### Vera Rubin GPUs to ship in the second half of 2026 - Claimant: Nvidia - Said: Mar 18, 2025 - Quote: "The centerpiece announcement was Vera Rubin, first teased at Computex 2024 and now scheduled for release in the second half of 2026." - Source: https://arstechnica.com/ai/2025/03/nvidia-announces-rubin-ultra-and-feynman-ai-chips-for-2027-and-2028/ - Resolution criteria (our interpretation): Our interpretation: Vera Rubin systems are commercially available from Nvidia partners by 31 Dec 2026. - Resolve by: 2026-12-31 - Verdict: Pending - URL: https://sipoly.si/claim/nvidia-vera-rubin-h2-2026/ ### An AI company will claim AGI "probably in 2026 or 2027" - Claimant: Kevin Roose - Said: Mar 14, 2025 - Quote: "I believe that very soon — probably in 2026 or 2027, but possibly as soon as this year — one or more A.I. companies will claim they’ve created an artificial general intelligence" - Source: https://www.nytimes.com/2025/03/14/technology/why-im-feeling-the-agi.html - Resolution criteria (our interpretation): Our interpretation: correct if by 31 Dec 2027 a major AI company formally and publicly claims to have created AGI. - Resolve by: 2027-12-31 - Verdict: Open - URL: https://sipoly.si/claim/roose-lab-claims-agi-2026-27/ ### Bet: AI produces Annals-quality number theory papers for under $100k each within 5 years - Claimant: Tamay Besiroglu - Said: Mar 11, 2025 - Quote: "I bet @littmath that in 5 years AI will be able to produce Annals-quality Number Theory papers at an inference budget at or below $100k/paper, at 3:1 odds in my favor." - Source: https://x.com/tamaybes/status/1899262088369106953 - Resolution criteria (our interpretation): Our interpretation: by 11 Mar 2030, AI produces number-theory papers judged of Annals of Mathematics quality (per the bet terms with Daniel Litt) at or below $100k inference cost per paper. - Resolve by: 2030-03-11 - Verdict: Open - URL: https://sipoly.si/claim/besiroglu-annals-number-theory-2030/ ### AI writing 90% of code in three to six months - Claimant: Dario Amodei - Said: Mar 10, 2025 - Quote: "I think we’ll be there in three to six months—where AI is writing 90 percent of the code." - Source: https://www.cfr.org/event/ceo-speaker-series-dario-amodei-anthropic - Resolution criteria (our interpretation): Our interpretation: by 10 Sep 2025, credible industry-wide measures show AI producing ~90% of code written by professional developers (not only at one company). - Resolve by: 2025-09-10 - Verdict: Wrong - URL: https://sipoly.si/claim/amodei-90pct-code/ ### More personalized Siri "in the coming year" - Claimant: Apple - Said: Mar 7, 2025 - Quote: "It’s going to take us longer than we thought to deliver on these features and we anticipate rolling them out in the coming year." - Source: https://daringfireball.net/2025/03/apple_is_delaying_the_more_personalized_siri_apple_intelligence_features - Resolution criteria (our interpretation): Our interpretation: personal-context Siri features ship to the public by 31 Mar 2026. - Resolve by: 2026-03-31 - Verdict: Broken - URL: https://sipoly.si/claim/apple-personal-siri-coming-year/ ### Powerful AI systems in late 2026 or early 2027 - Claimant: Anthropic - Said: Mar 6, 2025 - Quote: "we expect powerful AI systems will emerge in late 2026 or early 2027." - Source: https://www.anthropic.com/news/anthropic-s-recommendations-ostp-u-s-ai-action-plan - Resolution criteria (our interpretation): Our interpretation, using Anthropic's own definition: by 31 Mar 2027 a system exists matching or exceeding Nobel-level researchers across most listed disciplines. - Resolve by: 2027-03-31 - Verdict: Open - URL: https://sipoly.si/claim/anthropic-powerful-ai-2026-27/ ### GPT-4.5 in "weeks", GPT-5 in "months" - Claimant: OpenAI - Said: Feb 12, 2025 - Quote: "Altman didn’t say exactly when GPT-4.5 and GPT-5 might be released, but he did give a vague estimate of “weeks / months.”" - Source: https://www.theverge.com/news/611365/openai-gpt-4-5-roadmap-sam-altman-orion - Resolution criteria (our interpretation): Our interpretation: GPT-4.5 released within ~8 weeks and GPT-5 within 2025. - Resolve by: 2025-12-31 - Verdict: Kept - URL: https://sipoly.si/claim/openai-gpt45-gpt5-roadmap/ ### 75% chance of AGI (replacing the majority of human jobs) by end of 2028 - Claimant: Andrew Critch - Said: Feb 6, 2025 - Quote: "p(AGI by eoy 2028) = 75%" - Source: https://x.com/AndrewCritchPhD/status/1887541390097453100 - Resolution criteria (our interpretation): Our interpretation: AI that, at runtime and for less than a human costs, can replace humans in the power-weighted majority of jobs exists by 31 Dec 2028. - Resolve by: 2028-12-31 - Verdict: Open - URL: https://sipoly.si/claim/critch-agi-2028-75pct/ ### 40% chance of AGI (replacing the majority of human jobs) by end of 2026 - Claimant: Andrew Critch - Said: Feb 6, 2025 - Quote: "p(AGI by eoy 2026) = 40%" - Source: https://x.com/AndrewCritchPhD/status/1887541390097453100 - Resolution criteria (our interpretation): Our interpretation: AI that, at runtime and for less than a human costs, can replace humans in the power-weighted majority of jobs exists by 31 Dec 2026. - Resolve by: 2026-12-31 - Verdict: Open - URL: https://sipoly.si/claim/critch-agi-2026-40pct/ ### 20% chance of AGI (replacing the majority of human jobs) by end of 2025 - Claimant: Andrew Critch - Said: Feb 6, 2025 - Quote: "p(AGI by eoy 2025) = 20%" - Source: https://x.com/AndrewCritchPhD/status/1887541390097453100 - Resolution criteria (our interpretation): Our interpretation: AI that, at runtime and for less than a human costs, can replace humans in the power-weighted majority of jobs exists by 31 Dec 2025. - Resolve by: 2025-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/critch-agi-2025-20pct/ ### Within a decade AI makes great medical advice and tutoring "free, commonplace" - Claimant: Bill Gates - Said: Feb 2025 - Quote: "with AI, over the next decade, that will become free, commonplace — great medical advice, great tutoring" - Source: https://www.cnbc.com/2025/03/26/bill-gates-on-ai-humans-wont-be-needed-for-most-things.html - Resolution criteria (our interpretation): Our interpretation: by Feb 2035, AI medical advice and tutoring of a quality comparable to a great doctor or teacher are freely and widely available. - Resolve by: 2035-02-28 - Verdict: Open - URL: https://sipoly.si/claim/gates-humans-not-needed-decade/ ### At least six Indian developers will build foundation models within 8-10 months - Claimant: Ashwini Vaishnaw - Said: Jan 30, 2025 - Quote: "“We believe there are at least six developers who will be able to create foundation models in the next eight to 10 months on the outer limit and four to six months on an aggressive, optimistic estimate.”" - Source: https://www.business-standard.com/technology/tech-news/india-s-foundation-artificial-intelligence-model-in-8-10-months-vaishnaw-125013001688_1.html - Resolution criteria (our interpretation): Our interpretation: correct if by 30 Nov 2025 at least six Indian developers have released sovereign foundation models under or alongside the IndiaAI Mission. - Resolve by: 2025-11-30 - Verdict: Wrong - URL: https://sipoly.si/claim/vaishnaw-india-foundation-models-8-10m/ ### 2025 "may very well be" the year Llama and open source become the most advanced models - Claimant: Mark Zuckerberg - Said: Jan 29, 2025 - Quote: "I think this will very well be the year when Llama and open source become the most advanced and widely used AI models as well." - Source: https://www.fool.com/earnings/call-transcripts/2025/01/29/meta-platforms-meta-q4-2024-earnings-call-transcri/ - Resolution criteria (our interpretation): Our interpretation: at some point in 2025, a Llama model (or open-weight models generally) tops major independent capability leaderboards and is the most widely used model family. - Resolve by: 2025-12-31 - Verdict: Wrong - URL: https://sipoly.si/claim/zuckerberg-llama-most-advanced-2025/ ### Optimus "production design 2" to launch in 2026 - Claimant: Tesla - Said: Jan 29, 2025 - Quote: "Those Optimus in use at the Tesla factories for production design 1 will inform how will we change for production design 2, which we expect to launch next year." - Source: https://www.fool.com/earnings/call-transcripts/2025/01/29/tesla-tsla-q4-2024-earnings-call-transcript/ - Resolution criteria (our interpretation): Tesla starts production of a second Optimus production design during 2026. - Resolve by: 2026-12-31 - Verdict: Pending - URL: https://sipoly.si/claim/tesla-optimus-v2-2026/ ### "Several thousand" Optimus robots built in 2025 - Claimant: Tesla - Said: Jan 29, 2025 - Quote: "Will we succeed in building 10,000 exactly by the end of December this year? Probably not, but will we succeed in making several thousand? Yes, I think we will." - Source: https://www.fool.com/earnings/call-transcripts/2025/01/29/tesla-tsla-q4-2024-earnings-call-transcript/ - Resolution criteria (our interpretation): At least 2,000 Optimus robots built in calendar 2025. - Resolve by: 2025-12-31 - Verdict: Pending - URL: https://sipoly.si/claim/tesla-optimus-several-thousand-2025/ ### Unsupervised FSD paid service in Austin in June 2025 - Claimant: Tesla - Said: Jan 29, 2025 - Quote: "We’re going to be launching unsupervised full self-driving as a paid service in Austin in June." - Source: https://www.fool.com/earnings/call-transcripts/2025/01/29/tesla-tsla-q4-2024-earnings-call-transcript/ - Resolution criteria (our interpretation): Paid rides with no Tesla employee in the vehicle, open to the public in Austin, by 30 Jun 2025. - Resolve by: 2025-06-30 - Verdict: Partly kept - URL: https://sipoly.si/claim/tesla-austin-unsupervised-june/ ### AI Action Plan within 180 days (EO 14179) - Claimant: The White House (U.S. executive branch) - Said: Jan 23, 2025 - Quote: "Within 180 days of this order, the Assistant to the President for Science and Technology (APST), the Special Advisor for AI and Crypto, and the Assistant to the President for National Security Affairs (APNSA)" - Source: https://www.whitehouse.gov/presidential-actions/2025/01/removing-barriers-to-american-leadership-in-artificial-intelligence/ - Resolution criteria (our interpretation): Action Plan delivered within 180 days (by 22 Jul 2025). - Resolve by: 2025-07-22 - Verdict: Kept - URL: https://sipoly.si/claim/wh-eo14179-action-plan/ ### Current LLM paradigm has a shelf life of "probably three to five years" - Claimant: Yann LeCun - Said: Jan 23, 2025 - Quote: "I think the shelf life of the current [LLM] paradigm is fairly short, probably three to five years" - Source: https://techcrunch.com/2025/01/23/metas-yann-lecun-predicts-a-new-ai-architectures-paradigm-within-5-years-and-decade-of-robotics/ - Resolution criteria (our interpretation): Our interpretation: by 23 Jan 2030, LLMs are no longer the central component of leading AI systems. - Resolve by: 2030-01-23 - Verdict: Open - URL: https://sipoly.si/claim/lecun-llm-shelf-life/ ### Stargate: $500B of US AI infrastructure over four years - Claimant: OpenAI - Said: Jan 21, 2025 - Quote: "The Stargate Project is a new company which intends to invest $500 billion over the next four years building new AI infrastructure for OpenAI in the United States." - Source: https://openai.com/index/announcing-the-stargate-project/ - Resolution criteria (our interpretation): Our interpretation: credible disclosures show roughly $500B invested in Stargate US AI infrastructure by 21 Jan 2029. - Resolve by: 2029-01-21 - Verdict: Pending - URL: https://sipoly.si/claim/stargate-500b-four-years/ ### UK to publish a long-term AI infrastructure plan within 6 months, backed by a 10-year investment commitment - Claimant: UK Labour Party / UK Government - Said: Jan 13, 2025 - Quote: "Set out, within 6 months, a long-term plan for the UK’s AI infrastructure needs, backed by a 10-year investment commitment." - Source: https://www.gov.uk/government/publications/ai-opportunities-action-plan/ai-opportunities-action-plan - Resolution criteria (our interpretation): Our interpretation: correct if by mid-Jul 2025 the government publishes a long-term AI infrastructure plan with a multi-year (about 10-year) investment commitment; partly if the plan appears on time with a shorter funding horizon. - Resolve by: 2025-07-31 - Verdict: Partly kept - URL: https://sipoly.si/claim/uk-aiop-infrastructure-plan-6m/ ### UK to expand its AI Research Resource compute "by at least 20x by 2030" - Claimant: UK Labour Party / UK Government - Said: Jan 13, 2025 - Quote: "Expand the capacity of the AI Research Resource ( AIRR ) by at least 20x by 2030 - starting within 6 months." - Source: https://www.gov.uk/government/publications/ai-opportunities-action-plan/ai-opportunities-action-plan - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2030, UK government reports AIRR capacity at least 20 times its January 2025 level. - Resolve by: 2030-12-31 - Verdict: Pending - URL: https://sipoly.si/claim/uk-airr-20x-by-2030/ ### "A 50% chance of AGI in the next 3 years" (from Jan 2025) - Claimant: Shane Legg - Said: Jan 10, 2025 - Quote: "Of course this now means a 50% chance of AGI in the next 3 years!" - Source: https://x.com/ShaneLegg/status/1877726711027990738 - Resolution criteria (our interpretation): Our interpretation: by 10 Jan 2028, a system exists that Google DeepMind or broad expert consensus accepts as AGI (able to do the cognitive tasks people typically can). - Resolve by: 2028-01-10 - Verdict: Open - URL: https://sipoly.si/claim/legg-agi-50pct-by-2028/ ### "Probably in 2025" an AI that can be a mid-level engineer - Claimant: Mark Zuckerberg - Said: Jan 10, 2025 - Quote: "Probably in 2025, we at Meta, as well as the other companies that are basically working on this, are going to have an AI that can effectively be a sort of mid-level engineer that you have at your company that can write code" - Source: https://fortune.com/2025/01/24/mark-zuckerberg-ai-engineer-capex-spend/ - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2025, Meta or a peer deploys an AI agent doing the work of a mid-level software engineer end-to-end. - Resolve by: 2025-12-31 - Verdict: Open - URL: https://sipoly.si/claim/zuckerberg-midlevel-engineer-2025/ ### AGI "will probably get developed during this president's term" - Claimant: Sam Altman - Said: Jan 6, 2025 - Quote: "AGI will probably get developed during this president's term, and getting that right seems really important." - Source: https://www.businessinsider.com/sam-altman-hopes-donald-trump-build-new-ai-infrastructure-2025-1 - Resolution criteria (our interpretation): Our interpretation: AGI (as broadly recognized, not merely self-declared) is developed before 20 Jan 2029. - Resolve by: 2029-01-20 - Verdict: Open - URL: https://sipoly.si/claim/altman-agi-this-term/ ### AI agents may "join the workforce" in 2025 - Claimant: Sam Altman - Said: Jan 5, 2025 - Quote: "We believe that, in 2025, we may see the first AI agents “join the workforce” and materially change the output of companies." - Source: https://blog.samaltman.com/reflections - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2025, AI agents are deployed doing end-to-end work at companies with a measurable, reported effect on company output. - Resolve by: 2025-12-31 - Verdict: Unresolvable / too vague - URL: https://sipoly.si/claim/altman-2025-agents-workforce/ ### About $80B on AI datacenters in FY2025 - Claimant: Microsoft - Said: Jan 3, 2025 - Quote: "In FY 2025, Microsoft is on track to invest approximately $80 billion to build out AI-enabled datacenters" - Source: https://blogs.microsoft.com/on-the-issues/2025/01/03/the-golden-opportunity-for-american-ai/ - Resolution criteria (our interpretation): Microsoft-reported capital spending (incl. finance leases) for FY2025 (ended 30 Jun 2025) within ±15% of $80B. - Resolve by: 2025-06-30 - Verdict: Pending - URL: https://sipoly.si/claim/microsoft-80b-fy2025/ ### 2025 "could well be" the year AI company valuations start to fall - Claimant: Gary Marcus - Said: Jan 2025 - Quote: "2025 could well be the year in which valuations for major AI companies start to fall." - Source: https://garymarcus.substack.com/p/25-ai-predictions-for-2025-from-marcus - Resolution criteria (our interpretation): Our interpretation: correct if valuations of the leading private AI labs (OpenAI, Anthropic, xAI) were lower at end-2025 than at the start of 2025. - Resolve by: 2025-12-31 - Verdict: Wrong - URL: https://sipoly.si/claim/marcus-2025-valuations-may-fall/ ### Generative-AI copyright lawsuits will continue throughout 2025 - Claimant: Gary Marcus - Said: Jan 2025 - Quote: "Copyright lawsuits over generative AI will continue throughout the year." - Source: https://garymarcus.substack.com/p/25-ai-predictions-for-2025-from-marcus - Resolution criteria (our interpretation): Our interpretation: correct if major generative-AI copyright suits remain active or reach significant rulings or settlements during 2025. - Resolve by: 2025-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/marcus-2025-copyright-lawsuits-continue/ ### 30%: a major AI lab will formally claim AGI in 2025 - Claimant: Vox Future Perfect - Said: Jan 2025 - Quote: "A major lab will formally claim it has achieved AGI (30 percent)" - Source: https://www.vox.com/future-perfect/392241/2025-new-year-predictions-trump-musk-artificial-intelligence - Resolution criteria (our interpretation): Our interpretation: YES if in 2025 a major AI lab formally announces it has achieved AGI. - Resolve by: 2025-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/vox-2025-lab-claims-agi/ ### AI-designed drugs in clinical trials by the end of 2025 - Claimant: Isomorphic Labs - Said: 2025 - Quote: "Hassabis had said last year the company would have AI-designed drugs in clinical trials by the end of 2025." - Source: https://www.channelnewsasia.com/business/google-backed-isomorphic-labs-delays-clinical-trial-timeline-5871841 - Resolution criteria (our interpretation): Our interpretation: Isomorphic Labs doses the first human in a clinical trial of an AI-designed drug by 31 Dec 2025. - Resolve by: 2025-12-31 - Verdict: Broken - URL: https://sipoly.si/claim/isomorphic-trials-end-2025/ ### Hallucinations "will continue to haunt generative AI" in 2025 - Claimant: Gary Marcus - Said: Jan 1, 2025 - Quote: "Hallucinations (which should really be called confabulations) will continue to haunt generative AI." - Source: https://garymarcus.substack.com/p/25-ai-predictions-for-2025-from-marcus - Resolution criteria (our interpretation): Our interpretation: correct if at end of 2025 leading labs still acknowledge hallucinations as an unsolved problem in their frontier models. - Resolve by: 2025-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/marcus-2025-hallucinations-persist/ ### 2025: humanoid hype, but nothing "remotely as capable as Rosie the Robot" - Claimant: Gary Marcus - Said: Jan 1, 2025 - Quote: "Humanoid robotics will see a lot of hype, but nobody will release anything to remotely as capable as Rosie the Robot ." - Source: https://garymarcus.substack.com/p/25-ai-predictions-for-2025-from-marcus - Resolution criteria (our interpretation): Our interpretation: correct if by 31 Dec 2025 no general-purpose household humanoid robot (cooking, cleaning, varied chores without teleoperation) is released to customers. - Resolve by: 2025-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/marcus-2025-no-rosie-robot/ ### Truly driverless cars limited to a modest number of cities in 2025 - Claimant: Gary Marcus - Said: Jan 1, 2025 - Quote: "Truly driverless cars, in which no human is required to attend to traffic, will continue to see usage limited to a modest number of cities, mainly in the West, mainly in good weather." - Source: https://garymarcus.substack.com/p/25-ai-predictions-for-2025-from-marcus - Resolution criteria (our interpretation): Our interpretation: at end-2025, fully driverless public robotaxi service operates in no more than ~a dozen metro areas worldwide. - Resolve by: 2025-12-31 - Verdict: Open - URL: https://sipoly.si/claim/marcus-driverless-limited-2025/ ### AI agents hyped in 2025 but far from reliable - Claimant: Gary Marcus - Said: Jan 1, 2025 - Quote: "AI “Agents” will be endlessly hyped throughout 2025 but far from reliable, except possibly in very narrow use cases." - Source: https://garymarcus.substack.com/p/25-ai-predictions-for-2025-from-marcus - Resolution criteria (our interpretation): No operational threshold for "far from reliable". - Resolve by: 2025-12-31 - Verdict: Unresolvable / too vague - URL: https://sipoly.si/claim/marcus-agents-unreliable-2025/ ### Less than 10% of the workforce replaced by AI in 2025 - Claimant: Gary Marcus - Said: Jan 1, 2025 - Quote: "Less than 10% of the work force will be replaced by AI. Probably less than 5%." - Source: https://garymarcus.substack.com/p/25-ai-predictions-for-2025-from-marcus - Resolution criteria (our interpretation): Under 10% of the US workforce displaced by AI by 31 Dec 2025. - Resolve by: 2025-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/marcus-workforce-under-10pct/ ### Few if any radiologists replaced by AI in 2025 - Claimant: Gary Marcus - Said: Jan 1, 2025 - Quote: "Few if any radiologists will be replaced by AI (contra Hinton’s infamous 2016 prediction)." - Source: https://garymarcus.substack.com/p/25-ai-predictions-for-2025-from-marcus - Resolution criteria (our interpretation): No meaningful net reduction in employed radiologists attributable to AI in 2025. - Resolve by: 2025-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/marcus-radiologists-2025/ ### No system will solve more than 4 of the Marcus-Brundage tasks in 2025 - Claimant: Gary Marcus - Said: Jan 1, 2025 - Quote: "No single system will solve more than 4 of the AI 2027 Marcus-Brundage tasks by the end of 2025." - Source: https://garymarcus.substack.com/p/25-ai-predictions-for-2025-from-marcus - Resolution criteria (our interpretation): No single AI system reliably performs 5+ of the 10 tasks by 31 Dec 2025. - Resolve by: 2025-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/marcus-brundage-tasks-2025/ ### No AGI in 2025 - Claimant: Gary Marcus - Said: Jan 1, 2025 - Quote: "We will not see artificial general intelligence this year, despite claims by Elon Musk to the contrary." - Source: https://garymarcus.substack.com/p/25-ai-predictions-for-2025-from-marcus - Resolution criteria (our interpretation): No system broadly accepted as AGI by 31 Dec 2025. - Resolve by: 2025-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/marcus-no-agi-2025/ ### Bet: AI will do at least 8 of Marcus's 10 hard tasks by the end of 2027 - Claimant: Miles Brundage - Said: Dec 30, 2024 - Quote: "If there exist AI systems that can perform 8 of the 10 tasks below by the end of 2027, as determined by our panel of judges, Gary will donate $2,000 to a charity of Miles’ choice" - Source: https://garymarcus.substack.com/p/where-will-ai-be-at-the-end-of-2027 - Resolution criteria (our interpretation): Our interpretation: correct for Brundage if the agreed judging panel finds AI can do at least 8 of the 10 tasks by end-2027. - Resolve by: 2028-06-30 - Verdict: Open - URL: https://sipoly.si/claim/brundage-2027-bet-8-of-10-tasks/ ### Bet (10:1 odds): AI will not do 8 of 10 hard tasks by the end of 2027 - Claimant: Gary Marcus - Said: Dec 30, 2024 - Quote: "If there exist AI systems that can perform 8 of the 10 tasks below by the end of 2027, as determined by our panel of judges, Gary will donate $2,000 to a charity of Miles’ choice; if AI can do fewer than 8, Miles will donate $20,000 to a charity of Gary’s choice." - Source: https://garymarcus.substack.com/p/where-will-ai-be-at-the-end-of-2027 - Resolution criteria (our interpretation): Our interpretation: correct for Marcus if the agreed judging panel finds AI can do fewer than 8 of the 10 tasks by end-2027; resolved by the panel's ruling. - Resolve by: 2028-06-30 - Verdict: Open - URL: https://sipoly.si/claim/marcus-2027-bet-fewer-than-8-tasks/ ### AIs smarter than people "within probably the next 20 years" - Claimant: Geoffrey Hinton - Said: Dec 27, 2024 - Quote: "most of the experts in the field think that sometime, within probably the next 20 years, we’re going to develop AIs that are smarter than people." - Source: https://www.theguardian.com/technology/2024/dec/27/godfather-of-ai-raises-odds-of-the-technology-wiping-out-humanity-over-next-30-years - Resolution criteria (our interpretation): Our interpretation: by 27 Dec 2044, AI systems exceed human experts at most cognitive tasks. - Resolve by: 2044-12-27 - Verdict: Open - URL: https://sipoly.si/claim/hinton-smarter-than-people-20y/ ### "10% to 20%" chance AI leads to human extinction within three decades - Claimant: Geoffrey Hinton - Said: Dec 27, 2024 - Quote: "said there was a “10% to 20%” chance that AI would lead to human extinction within the next three decades." - Source: https://www.theguardian.com/technology/2024/dec/27/godfather-of-ai-raises-odds-of-the-technology-wiping-out-humanity-over-next-30-years - Resolution criteria (our interpretation): Our interpretation: a probability statement over 30 years; it cannot be scored on a single outcome before 2054. Kept for the record. - Resolve by: 2054-12-27 - Verdict: Open - URL: https://sipoly.si/claim/hinton-extinction-10-20pct-30y/ ### 60%: an OpenAI o-series model solves a Millennium Prize problem in 2025 - Claimant: Aidan McLaughlin - Said: Dec 21, 2024 - Quote: "i think it’s likely (p=.6) that an o-series model solves a millennium prize math problem in 2025" - Source: https://x.com/aidan_mclau/status/1870462987842236910 - Resolution criteria (our interpretation): Our interpretation: an OpenAI o-series model produces an accepted solution to one of the Clay Millennium Prize problems during 2025. - Resolve by: 2025-12-31 - Verdict: Wrong - URL: https://sipoly.si/claim/mclaughlin-millennium-prize-2025/ ### o3-mini "around the end of January" - Claimant: OpenAI - Said: Dec 20, 2024 - Quote: "Altman said the company plans to launch that model "around the end of January," with o3 following "shortly after that."" - Source: https://www.engadget.com/ai/openais-next-generation-o3-model-will-arrive-early-next-year-191707632.html - Resolution criteria (our interpretation): o3-mini publicly available by 7 Feb 2025. - Resolve by: 2025-02-07 - Verdict: Kept - URL: https://sipoly.si/claim/openai-o3-mini-january/ ### Apple Intelligence on iPhone and iPad in the EU starting April 2025 - Claimant: Apple - Said: Dec 11, 2024 - Quote: "This April, Apple Intelligence features will start to roll out to iPhone and iPad users in the EU." - Source: https://www.apple.com/newsroom/2024/12/apple-intelligence-now-features-image-playground-genmoji-and-more/ - Resolution criteria (our interpretation): Our interpretation: iPhone and iPad users in the EU can use Apple Intelligence features by 30 Apr 2025. - Resolve by: 2025-04-30 - Verdict: Kept - URL: https://sipoly.si/claim/apple-intelligence-eu-iphone-april-2025/ ### Apple Intelligence in more languages (Chinese, French, German, Japanese…) starting in April 2025 - Claimant: Apple - Said: Dec 11, 2024 - Quote: "Additional languages, including Chinese, English (India), English (Singapore), French, German, Italian, Japanese, Korean, Portuguese, Spanish, and Vietnamese will be coming throughout the year, with an initial set arriving in a software update in April." - Source: https://www.apple.com/newsroom/2024/12/apple-intelligence-now-features-image-playground-genmoji-and-more/ - Resolution criteria (our interpretation): Our interpretation: a software update in (or before the end of) April 2025 adds Apple Intelligence support for an initial set of the listed languages. - Resolve by: 2025-04-30 - Verdict: Kept - URL: https://sipoly.si/claim/apple-intelligence-languages-april-2025/ ### Waymo One open to riders in Miami in 2026 - Claimant: Waymo - Said: Dec 5, 2024 - Quote: "Through our new fleet partnership with Moove , a global leader in innovative mobility solutions, we’ll work to open our doors to riders in 2026, offering our ride-hailing service via the Waymo One app." - Source: https://waymo.com/blog/2024/12/next-stop-miami - Resolution criteria (our interpretation): Our interpretation: fully autonomous Waymo rides are available to public riders in Miami during 2026. - Resolve by: 2026-12-31 - Verdict: Kept - URL: https://sipoly.si/claim/waymo-miami-2026/ ### Commercial driverless trucking in Texas in April 2025 - Claimant: Aurora Innovation - Said: Oct 30, 2024 - Quote: "With additional visibility on the time needed to complete the aforementioned remaining validation, we now expect to launch commercially in April 2025." - Source: https://ir.aurora.tech/sec-filings/all-sec-filings/content/0001828108-24-000138/aurora-3q24xshareholder_.htm - Resolution criteria (our interpretation): Our interpretation: Aurora begins commercial driverless freight service by about the end of April 2025 (we allow the first week of May). - Resolve by: 2025-05-07 - Verdict: Kept - URL: https://sipoly.si/claim/aurora-driverless-trucks-april-2025/ ### AGI "likely to emerge in the next four years" (from Oct 2024) - Claimant: Thomas L. Friedman - Said: Oct 29, 2024 - Quote: "the birth of artificial general intelligence, or A.G.I., which is likely to emerge in the next four years" - Source: https://www.nytimes.com/2024/10/29/opinion/artificial-intelligence-harris-trump-election.html - Resolution criteria (our interpretation): Our interpretation: correct if by 29 Oct 2028 an AI system is broadly recognized by independent experts as AGI (matching or beating humans across most cognitive work); a developer's own claim does not count on its own. - Resolve by: 2028-10-29 - Verdict: Open - URL: https://sipoly.si/claim/friedman-agi-next-four-years/ ### Instinct MI350 accelerators to launch in the second half of 2025 - Claimant: AMD - Said: Oct 10, 2024 - Quote: "AMD also shared new details on next-gen AMD Instinct MI350 series accelerators expected to launch in the second half of 2025" - Source: https://ir.amd.com/news-events/press-releases/detail/1218/amd-unveils-leadership-ai-solutions-at-advancing-ai-2024 - Resolution criteria (our interpretation): Our interpretation: AMD launches MI350-series accelerators (announced as shipping or available) by 31 Dec 2025. - Resolve by: 2025-12-31 - Verdict: Kept - URL: https://sipoly.si/claim/amd-mi350-h2-2025/ ### Data-center AI accelerator market to reach $500 billion by 2028 - Claimant: AMD - Said: Oct 10, 2024 - Quote: "Looking ahead, we see the data center AI accelerator market growing to $500 billion by 2028." - Source: https://ir.amd.com/news-events/press-releases/detail/1218/amd-unveils-leadership-ai-solutions-at-advancing-ai-2024 - Resolution criteria (our interpretation): Our interpretation: credible market estimates (e.g. from AMD, Nvidia filings, analysts) put 2028 data-center AI accelerator revenue at $500B or more. - Resolve by: 2029-06-30 - Verdict: Open - URL: https://sipoly.si/claim/amd-ai-accelerator-tam-500b-2028/ ### AI systems improving their own code in ~5 years; 80-90% of expert ability in every field in 6-8 years - Claimant: Eric Schmidt - Said: Oct 2024 - Quote: "In the industry it is believed that somewhere around five years, no one knows exactly, the systems will begin to be able to write their own code, that is, they literally will take their code and make it better." - Source: https://www.youtube.com/watch?v=cfbD9bsPlFQ - Resolution criteria (our interpretation): Our interpretation: correct if by 31 Oct 2029 frontier AI systems are documented (by a lab or independent evaluators) as substantially improving their own code or training process, i.e. recursive self-improvement in practice. - Resolve by: 2029-10-31 - Verdict: Open - URL: https://sipoly.si/claim/schmidt-recursive-code-5y-experts-6-8y/ ### Powerful AI "could come as early as 2026" - Claimant: Dario Amodei - Said: Oct 2024 - Quote: "I think it could come as early as 2026, though there are also ways it could take much longer." - Source: https://www.darioamodei.com/essay/machines-of-loving-grace - Resolution criteria (our interpretation): Our interpretation: a system meeting the essay's "powerful AI" definition exists by 31 Dec 2026. - Resolve by: 2026-12-31 - Verdict: Open - URL: https://sipoly.si/claim/amodei-powerful-ai-2026/ ### Superintelligence possible "in a few thousand days" - Claimant: Sam Altman - Said: Sep 23, 2024 - Quote: "It is possible that we will have superintelligence in a few thousand days (!); it may take longer, but I’m confident we’ll get there." - Source: https://ia.samaltman.com/ - Resolution criteria (our interpretation): No operational definition or date given ("possible", "may take longer"). - Resolve by: 2032-12-31 - Verdict: Unresolvable / too vague - URL: https://sipoly.si/claim/altman-superintelligence-thousand-days/ ### Waymo on Uber in Austin and Atlanta "beginning in early 2025" - Claimant: Waymo - Said: Sep 13, 2024 - Quote: "Today, Waymo and Uber are announcing an expanded partnership to bring the Waymo One experience to Austin and Atlanta, only on the Uber app, beginning in early 2025." - Source: https://waymo.com/blog/2024/09/waymo-and-uber-expand-partnership - Resolution criteria (our interpretation): Our interpretation: public robotaxi rides on Uber in both Austin and Atlanta by the end of Q1 2025 ("early 2025"); partly if only one city launches in Q1 and the other follows within 2025. - Resolve by: 2025-03-31 - Verdict: Partly kept - URL: https://sipoly.si/claim/waymo-uber-austin-atlanta-early-2025/ ### "One billion agents with Agentforce by the end of 2025" - Claimant: Salesforce - Said: Sep 12, 2024 - Quote: "Our vision is bold: to empower one billion agents with Agentforce by the end of 2025." - Source: https://www.salesforce.com/news/press-releases/2024/09/12/agentforce-announcement/ - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2025, Salesforce reports one billion Agentforce agents deployed. - Resolve by: 2025-12-31 - Verdict: Broken - URL: https://sipoly.si/claim/salesforce-billion-agents-2025/ ### At least 30% of generative AI projects abandoned after proof of concept by end of 2025 - Claimant: Gartner - Said: Jul 29, 2024 - Quote: "At least 30% of generative AI (GenAI) projects will be abandoned after proof of concept by the end of 2025" - Source: https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025 - Resolution criteria (our interpretation): Our interpretation: credible enterprise surveys covering 2025 show at least 30% of generative-AI projects stopped after the proof-of-concept stage. - Resolve by: 2025-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/gartner-genai-30pct-abandoned-2025/ ### GPAI rules apply from 2 August 2025 - Claimant: European Union (AI Act) - Said: Jul 12, 2024 - Quote: "Chapter III Section 4, Chapter V, Chapter VII and Chapter XII and Article 78 shall apply from 2 August 2025" - Source: https://artificialintelligenceact.eu/article/113/ - Resolution criteria (our interpretation): GPAI obligations legally apply from 2 Aug 2025 without postponement. - Resolve by: 2025-08-02 - Verdict: Kept - URL: https://sipoly.si/claim/eu-gpai-obligations-aug-2025/ ### GPAI codes of practice ready "at the latest by 2 May 2025" - Claimant: European Union (AI Act) - Said: Jul 12, 2024 - Quote: "Codes of practice shall be ready at the latest by 2 May 2025." - Source: https://artificialintelligenceact.eu/article/56/ - Resolution criteria (our interpretation): GPAI code of practice finalised by 2 May 2025. - Resolve by: 2025-05-02 - Verdict: Partly kept - URL: https://sipoly.si/claim/eu-gpai-codes-may-2025/ ### Grok 3 by the end of 2024, after training on 100k H100s - Claimant: xAI - Said: Jul 1, 2024 - Quote: "Grok 3 end of year after training on 100k H100s should be really something special" - Source: https://x.com/elonmusk/status/1807643760584708363 - Resolution criteria (our interpretation): Our interpretation: xAI publicly releases Grok 3 by 31 Dec 2024. - Resolve by: 2024-12-31 - Verdict: Broken - URL: https://sipoly.si/claim/xai-grok3-end-2024/ ### "We are going to expand intelligence a millionfold by 2045" - Claimant: Ray Kurzweil - Said: Jun 29, 2024 - Quote: "We are going to expand intelligence a millionfold by 2045 and it is going to deepen our awareness and consciousness." - Source: https://www.theguardian.com/technology/article/2024/jun/29/ray-kurzweil-google-ai-the-singularity-is-nearer - Resolution criteria (our interpretation): Our interpretation: by 2045, humans routinely augment their intelligence through brain-computer interfaces connected to AI, and effective (human plus machine) intelligence is many orders of magnitude above 2024 levels. - Resolve by: 2045-12-31 - Verdict: Open - URL: https://sipoly.si/claim/kurzweil-millionfold-2045/ ### LLM hallucinations "much less of a problem, certainly by 2029" - Claimant: Ray Kurzweil - Said: Jun 29, 2024 - Quote: "LLM hallucinations [where they create nonsensical or inaccurate outputs] will become much less of a problem, certainly by 2029" - Source: https://www.theguardian.com/technology/article/2024/jun/29/ray-kurzweil-google-ai-the-singularity-is-nearer - Resolution criteria (our interpretation): Our interpretation: by end-2029, frontier models' hallucination rates on standard factuality benchmarks (e.g. SimpleQA-style or OpenAI/Vectara hallucination leaderboards) are a small fraction (≤ one-fifth) of 2024 levels, and hallucination is no longer named a top barrier to deployment. - Resolve by: 2029-12-31 - Verdict: Open - URL: https://sipoly.si/claim/kurzweil-hallucinations-solved-2029/ ### Human-level intelligence and AGI by 2029 - Claimant: Ray Kurzweil - Said: Jun 29, 2024 - Quote: "So 2029, both for human-level intelligence and for artificial general intelligence (AGI) – which is a little bit different." - Source: https://www.theguardian.com/technology/article/2024/jun/29/ray-kurzweil-google-ai-the-singularity-is-nearer - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2029, AI matches the most skilled humans in most domains. - Resolve by: 2029-12-31 - Verdict: Open - URL: https://sipoly.si/claim/kurzweil-2029/ ### AI "10,000 times smarter than humans" within 10 years - Claimant: Masayoshi Son - Said: Jun 21, 2024 - Quote: "Artificial intelligence that is 10,000 times smarter than humans will be here in 10 years, SoftBank CEO Masayoshi Son said on Friday in a rare public appearance." - Source: https://www.cnbc.com/2024/06/21/softbank-ceo-predicts-ai-that-is-10000-times-smarter-than-humans-.html - Resolution criteria (our interpretation): Our interpretation: by 21 Jun 2034, artificial superintelligence vastly exceeding the best humans at essentially all cognitive tasks exists. "10,000 times" is not measurable, so we score the ASI part. - Resolve by: 2034-06-21 - Verdict: Open - URL: https://sipoly.si/claim/son-asi-10000x-10-years/ ### AGI "one to 10 times smarter than humans" in the next three to five years - Claimant: Masayoshi Son - Said: Jun 21, 2024 - Quote: "Son said this tech is likely to be one to 10 times smarter than humans and will arrive in the next three-to-five years, earlier than he had anticipated." - Source: https://www.cnbc.com/2024/06/21/softbank-ceo-predicts-ai-that-is-10000-times-smarter-than-humans-.html - Resolution criteria (our interpretation): Our interpretation: AGI (AI at least as capable as humans across most cognitive tasks) exists by 21 Jun 2029. - Resolve by: 2029-06-21 - Verdict: Open - URL: https://sipoly.si/claim/son-agi-3-5-years/ ### Claude 3.5 Haiku and 3.5 Opus "later this year" - Claimant: Anthropic - Said: Jun 21, 2024 - Quote: "To complete the Claude 3.5 model family, we’ll be releasing Claude 3.5 Haiku and Claude 3.5 Opus later this year." - Source: https://www.anthropic.com/news/claude-3-5-sonnet - Resolution criteria (our interpretation): Both models released by 31 Dec 2024. - Resolve by: 2024-12-31 - Verdict: Partly kept - URL: https://sipoly.si/claim/anthropic-35-haiku-opus/ ### Binding regulation on the most powerful AI model developers - Claimant: UK Labour Party / UK Government - Said: Jun 13, 2024 - Quote: "Labour will ensure the safe development and use of AI models by introducing binding regulation on the handful of companies developing the most powerful AI models" - Source: https://labour.org.uk/change/kickstart-economic-growth/ - Resolution criteria (our interpretation): Our interpretation: binding legislation applying to frontier AI developers enacted in the UK during this Parliament (we use 4 Jul 2029). - Resolve by: 2029-07-04 - Verdict: Pending - URL: https://sipoly.si/claim/uk-labour-binding-ai-regulation/ ### In five years there will be more software engineers than today, not fewer - Claimant: François Chollet - Said: Jun 11, 2024 - Quote: "Sure. In five years, there will be more software engineers than there are today, not fewer." - Source: https://www.dwarkesh.com/p/francois-chollet - Resolution criteria (our interpretation): Our interpretation: correct if US employment of software developers/engineers (BLS OES/CPS) in mid-2029 is above its 2024 level. - Resolve by: 2029-06-11 - Verdict: Open - URL: https://sipoly.si/claim/chollet-more-swes-in-5y/ ### ChatGPT integration in iOS 18 and macOS Sequoia "later this year" (2024) - Claimant: Apple - Said: Jun 10, 2024 - Quote: "ChatGPT will come to iOS 18, iPadOS 18, and macOS Sequoia later this year, powered by GPT-4o." - Source: https://www.apple.com/newsroom/2024/06/introducing-apple-intelligence-for-iphone-ipad-and-mac/ - Resolution criteria (our interpretation): Our interpretation: ChatGPT integration ships to the public in iOS/iPadOS 18 and macOS Sequoia by 31 Dec 2024. - Resolve by: 2024-12-31 - Verdict: Kept - URL: https://sipoly.si/claim/apple-chatgpt-ios18-2024/ ### Individual training clusters costing hundreds of billions of dollars by 2028 - Claimant: Leopold Aschenbrenner - Said: Jun 4, 2024 - Quote: "We’re on the path to individual training clusters costing $100s of billions by 2028—clusters requiring power equivalent to a small/medium US state and more expensive than the International Space Station." - Source: https://situational-awareness.ai/racing-to-the-trillion-dollar-cluster/ - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2028, at least one single training cluster (one campus) costs $100B or more, per credible reporting or trackers. - Resolve by: 2028-12-31 - Verdict: Open - URL: https://sipoly.si/claim/aschenbrenner-100b-clusters-2028/ ### "Strikingly plausible" that models do the work of an AI researcher/engineer by 2027 - Claimant: Leopold Aschenbrenner - Said: Jun 4, 2024 - Quote: "I make the following claim: it is strikingly plausible that by 2027, models will be able to do the work of an AI researcher/engineer." - Source: https://situational-awareness.ai/from-gpt-4-to-agi/ - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2027, AI systems can do the job of a frontier-lab AI researcher/engineer (not just assist), per lab statements and independent evaluations. - Resolve by: 2027-12-31 - Verdict: Open - URL: https://sipoly.si/claim/aschenbrenner-ai-researcher-2027/ ### PhD-level AI for specific tasks in "a year and a half" - Claimant: Mira Murati - Said: Jun 2024 - Quote: "And then in the next couple of years, we're looking at PhD-level intelligence for specific tasks." - Source: https://www.youtube.com/watch?v=yUoj9B8OpR8 - Resolution criteria (our interpretation): Our interpretation: by Dec 2025, a widely available AI system performs at PhD level on specific expert tasks (e.g. matches PhD experts on GPQA-style questions in their field). - Resolve by: 2025-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/murati-phd-level-18-months/ ### AI will displace social drinking within five years - Claimant: Jonathan Ross - Said: May 29, 2024 - Quote: "Prediction: AI will displace social drinking within 5 years" - Source: https://x.com/JonathanRoss321/status/1795722941990314240 - Resolution criteria (our interpretation): Our interpretation: by 29 May 2029, credible data show AI social assistants (e.g. earbuds) have measurably reduced social alcohol consumption, e.g. surveys attributing reduced drinking to AI use. - Resolve by: 2029-05-29 - Verdict: Open - URL: https://sipoly.si/claim/ross-ai-displaces-social-drinking/ ### Seoul commitment: publish a safety framework before the France summit - Claimant: Meta - Said: May 21, 2024 - Quote: "to demonstrate how they have achieved this by publishing a safety framework focused on severe risks by the upcoming AI Summit in France." - Source: https://www.gov.uk/government/publications/frontier-ai-safety-commitments-ai-seoul-summit-2024/frontier-ai-safety-commitments-ai-seoul-summit-2024 - Resolution criteria (our interpretation): A published severe-risk safety framework before the Paris AI Action Summit (10 Feb 2025). - Resolve by: 2025-02-10 - Verdict: Kept - URL: https://sipoly.si/claim/meta-seoul-framework/ ### Seoul commitment: publish a safety framework before the France summit - Claimant: Google / Google DeepMind - Said: May 21, 2024 - Quote: "to demonstrate how they have achieved this by publishing a safety framework focused on severe risks by the upcoming AI Summit in France." - Source: https://www.gov.uk/government/publications/frontier-ai-safety-commitments-ai-seoul-summit-2024/frontier-ai-safety-commitments-ai-seoul-summit-2024 - Resolution criteria (our interpretation): A published severe-risk safety framework before the Paris AI Action Summit (10 Feb 2025). - Resolve by: 2025-02-10 - Verdict: Kept - URL: https://sipoly.si/claim/google-seoul-framework/ ### Recall on Copilot+ PCs in preview starting June 18, 2024 - Claimant: Microsoft - Said: May 20, 2024 - Quote: "Now with Recall, in preview starting June 18, you can access virtually what you have seen or done on your PC in a way that feels like having photographic memory." - Source: https://blogs.microsoft.com/blog/2024/05/20/introducing-copilot-pcs/ - Resolution criteria (our interpretation): Our interpretation: Recall ships in preview on Copilot+ PCs (to general buyers, not only Insiders) on or about 18 Jun 2024. - Resolve by: 2024-06-18 - Verdict: Broken - URL: https://sipoly.si/claim/microsoft-recall-preview-june-2024/ ### Frontier Safety Framework "fully implemented by early 2025" - Claimant: Google / Google DeepMind - Said: May 17, 2024 - Quote: "We aim to have this initial framework fully implemented by early 2025." - Source: https://deepmind.google/discover/blog/introducing-the-frontier-safety-framework/ - Resolution criteria (our interpretation): Our interpretation: by 30 Apr 2025 Google DeepMind states the framework is applied in its model evaluation and governance. - Resolve by: 2025-04-30 - Verdict: Kept - URL: https://sipoly.si/claim/gdm-fsf-implemented-2025/ ### AI Overviews to reach over a billion people by the end of 2024 - Claimant: Google / Google DeepMind - Said: May 14, 2024 - Quote: "That means that this week, hundreds of millions of users will have access to AI Overviews, and we expect to bring them to over a billion people by the end of the year." - Source: https://blog.google/products/search/generative-ai-google-search-may-2024/ - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2024, Google says AI Overviews reach more than 1 billion users. - Resolve by: 2024-12-31 - Verdict: Kept - URL: https://sipoly.si/claim/google-ai-overviews-1b-2024/ ### Release the AlphaFold 3 model, including weights, for academic use within six months - Claimant: Google / Google DeepMind - Said: May 2024 - Quote: "posted on the social-media platform X that the team is “working on releasing the AF3 model (incl weights) for academic use” within six months." - Source: https://www.nature.com/articles/d41586-024-01463-0 - Resolution criteria (our interpretation): Our interpretation: by mid-Nov 2024, Google DeepMind makes AlphaFold 3 code plus model weights available for academic (non-commercial) use. - Resolve by: 2024-11-30 - Verdict: Kept - URL: https://sipoly.si/claim/google-deepmind-af3-code-6-months/ ### New GPT-4o Voice Mode alpha "in the coming weeks" - Claimant: OpenAI - Said: May 13, 2024 - Quote: "We’ll roll out a new version of Voice Mode with GPT‑4o in alpha within ChatGPT Plus in the coming weeks." - Source: https://openai.com/index/hello-gpt-4o/ - Resolution criteria (our interpretation): Our interpretation: alpha available to some ChatGPT Plus users within ~6 weeks (by 30 Jun 2024). - Resolve by: 2024-06-30 - Verdict: Partly kept - URL: https://sipoly.si/claim/openai-voice-mode-weeks/ ### Commercial driverless trucking launch by the end of 2024 - Claimant: Aurora Innovation - Said: May 2024 - Quote: "We are off to a strong start in 2024, driving purposefully toward our planned Commercial Launch at the end of the year and the subsequent scaling of our business." - Source: https://ir.aurora.tech/_assets/_e6f8f36cc43aa97c236ea5b9645f2ad2/aurora/db/956/9850/shareholder_letter/1Q24%2BShareholder%2BLetter.pdf - Resolution criteria (our interpretation): Our interpretation: Aurora runs commercial driverless (no one on board) freight hauls on public roads by 31 Dec 2024. - Resolve by: 2024-12-31 - Verdict: Broken - URL: https://sipoly.si/claim/aurora-driverless-trucks-end-2024/ ### Over $500 million of Gaudi AI accelerator revenue in 2024 - Claimant: Intel - Said: Apr 25, 2024 - Quote: "We now expect over $500 million in accelerated revenue in second half of 2024 and with increasing momentum into 2025 based on Gaudi 3's vastly superior TCO as well as our own expanding supply." - Source: https://www.fool.com/earnings/call-transcripts/2024/04/25/intel-intc-q1-2024-earnings-call-transcript/ - Resolution criteria (our interpretation): Our interpretation: Intel books more than $500M of Gaudi accelerator revenue in the second half of 2024 (on the same call the CFO said "greater than $500 million for the year"). - Resolve by: 2024-12-31 - Verdict: Broken - URL: https://sipoly.si/claim/intel-gaudi-500m-2024/ ### Llama 3: multimodal, multilingual, longer-context models "over the coming months" - Claimant: Meta - Said: Apr 18, 2024 - Quote: "Over the coming months, we’ll release multiple models with new capabilities including multimodality, the ability to converse in multiple languages, a much longer context window, and stronger overall capabilities." - Source: https://ai.meta.com/blog/meta-llama-3/ - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2024 Meta releases Llama 3-family models covering multilinguality, a much longer context window and multimodality. - Resolve by: 2024-12-31 - Verdict: Kept - URL: https://sipoly.si/claim/meta-llama3-coming-months/ ### AGI, defined as smarter than the smartest human, "within two years" - Claimant: Elon Musk - Said: Apr 8, 2024 - Quote: "If you define AGI as smarter than the smartest human, I think it’s probably next year, within two years" - Source: https://arstechnica.com/information-technology/2024/04/elon-musk-ai-will-be-smarter-than-any-human-around-the-end-of-next-year/ - Resolution criteria (our interpretation): Our interpretation: by 8 Apr 2026, an AI is smarter than the smartest human across domains. - Resolve by: 2026-04-08 - Verdict: Open - URL: https://sipoly.si/claim/musk-agi-within-two-years/ ### AI smarter than any one human "around the end of next year" - Claimant: Elon Musk - Said: Apr 8, 2024 - Quote: "My guess is we’ll have AI smarter than any one human probably around the end of next year" - Source: https://arstechnica.com/information-technology/2024/04/elon-musk-ai-will-be-smarter-than-any-human-around-the-end-of-next-year/ - Resolution criteria (our interpretation): Our interpretation: by ~31 Dec 2025 (grace to 31 Mar 2026), an AI system outperforms the best human in essentially every cognitive domain. - Resolve by: 2026-03-31 - Verdict: Wrong - URL: https://sipoly.si/claim/musk-ai-smarter-end-2025/ ### AI will add less than 0.55% to total factor productivity over 10 years - Claimant: Daron Acemoglu - Said: Apr 5, 2024 - Quote: "Consequently, predicted TFP gains over the next 10 years are even more modest and are predicted to be less than 0.55%." - Source: https://economics.mit.edu/sites/default/files/2024-04/The%20Simple%20Macroeconomics%20of%20AI.pdf - Resolution criteria (our interpretation): Our interpretation: by 2034, mainstream estimates (e.g. BLS/Fed/OECD decompositions) attribute less than 0.55% cumulative US TFP gain to AI over 2024-2034. - Resolve by: 2034-12-31 - Verdict: Open - URL: https://sipoly.si/claim/acemoglu-tfp-under-055-10y/ ### AI will do well on any specified test set ("AGI") within 5 years - Claimant: Jensen Huang - Said: Mar 19, 2024 - Quote: "If we specified AGI to be something very specific, a set of tests where a software program can do very well — or maybe 8% better than most people — I believe we will get there within 5 years" - Source: https://techcrunch.com/2024/03/19/agi-and-hallucinations/ - Resolution criteria (our interpretation): Our interpretation: by 19 Mar 2029, AI systems score better than most humans on essentially any standardized human test (professional, academic and logic exams). - Resolve by: 2029-03-19 - Verdict: Open - URL: https://sipoly.si/claim/huang-agi-tests-5-years/ ### Claims of AGI by 2030 are "laughable" - Claimant: Christopher Manning - Said: Mar 14, 2024 - Quote: "I do not believe human-level AI (artificial superintelligence, or the commonest sense of #AGI) is close at hand. AI has made breakthroughs, but the claim of AGI by 2030 is as laughable as claims of AGI by 1980 are in retrospect." - Source: https://x.com/chrmanning/status/1768291975005196326 - Resolution criteria (our interpretation): Our interpretation: correct if by 31 Dec 2030 no AI system is broadly accepted as human-level AGI. - Resolve by: 2030-12-31 - Verdict: Open - URL: https://sipoly.si/claim/manning-agi-2030-laughable/ ### By 2029, AI "probably smarter than all humans combined" - Claimant: Elon Musk - Said: Mar 13, 2024 - Quote: "By 2029, AI is probably smarter than all humans combined." - Source: https://x.com/elonmusk/status/1767738797276451090 - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2029, a single AI system (or coordinated AI collective) out-thinks humanity's combined intellectual output, e.g. it can do essentially all cognitive work faster and better than all human experts together. - Resolve by: 2029-12-31 - Verdict: Open - URL: https://sipoly.si/claim/musk-2029-smarter-than-all-humans/ ### "At least 3-5 years away from automating software engineering" (from Mar 2024) - Claimant: Bindu Reddy - Said: Mar 12, 2024 - Quote: "We are at least 3-5 years away from automating software engineering." - Source: https://x.com/bindureddy/status/1767608523859661120 - Resolution criteria (our interpretation): Our interpretation: correct if software engineering is not automated (AI doing the end-to-end work of most professional software engineers) before 12 Mar 2027. - Resolve by: 2027-03-12 - Verdict: Open - URL: https://sipoly.si/claim/reddy-swe-automation-3-5y/ ### "This week, @xAI will open source Grok" - Claimant: xAI - Said: Mar 11, 2024 - Quote: "This week, @xAI will open source Grok" - Source: https://siliconangle.com/2024/03/11/elon-musks-xai-open-source-grok-language-model/ - Resolution criteria (our interpretation): Grok model weights released openly by 17 Mar 2024. - Resolve by: 2024-03-17 - Verdict: Kept - URL: https://sipoly.si/claim/xai-open-source-grok/ ### By end of 2024: modest lasting corporate adoption and modest profits split 7-10 ways - Claimant: Gary Marcus - Said: Mar 10, 2024 - Quote: "Modest lasting corporate adoption • Modest profits, split 7-10 ways" - Source: https://x.com/GaryMarcus/status/1766871625075409381 - Resolution criteria (our interpretation): Our interpretation: correct if in 2024 enterprise adoption of generative AI stayed modest (mostly pilots) and model-maker profits were small or negative and spread across many firms. - Resolve by: 2024-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/marcus-2024-modest-adoption-profits/ ### By end of 2024: 7-10 GPT-4-level models - Claimant: Gary Marcus - Said: Mar 10, 2024 - Quote: "Prediction: By end of 2024 we will see • 7-10 GPT-4 level models" - Source: https://x.com/GaryMarcus/status/1766871625075409381 - Resolution criteria (our interpretation): Our interpretation: correct if by 31 Dec 2024 roughly 7-10 distinct models from different developers perform at about GPT-4 level, i.e. a crowded plateau rather than one leader far ahead. - Resolve by: 2024-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/marcus-2024-gpt4-level-models/ ### By end of 2024: no robust solution to hallucinations - Claimant: Gary Marcus - Said: Mar 10, 2024 - Quote: "No robust solution to hallucinations" - Source: https://x.com/GaryMarcus/status/1766871625075409381 - Resolution criteria (our interpretation): Our interpretation: correct if by 31 Dec 2024 frontier LLMs still hallucinate at rates that prevent unsupervised use in high-stakes settings and no lab claims a general fix. - Resolve by: 2024-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/marcus-2024-no-robust-hallucination-fix/ ### Human-level AGI most likely in 2029-2030, possibly as early as 2027 - Claimant: Ben Goertzel - Said: Mar 2024 - Quote: "Goertzel suggested 2029 or 2030 could be the likeliest years when humanity will build the first AGI agent, but that it could happen as early as 2027." - Source: https://www.livescience.com/technology/artificial-intelligence/ai-agi-singularity-in-2027-artificial-super-intelligence-sooner-than-we-think-ben-goertzel - Resolution criteria (our interpretation): Our interpretation: correct if a human-level AGI agent (broadly as capable as humans across cognitive tasks) is built by end-2030. - Resolve by: 2030-12-31 - Verdict: Open - URL: https://sipoly.si/claim/goertzel-agi-2027-2030/ ### About a 50% chance of AGI within 3 years (Feb 2024) - Claimant: Daniel Kokotajlo - Said: Feb 19, 2024 - Quote: "I expect to need the money sometime in the next 3 years, because that’s about when we get to 50% chance of AGI." - Source: https://www.greaterwrong.com/posts/CcqaJFf7TvAjuZFCx/retirement-accounts-and-short-timelines/comment/9tyDD2bqk4BH6Y7XX - Resolution criteria (our interpretation): Our interpretation: probabilistic (50%). YES if by 19 Feb 2027 an AI system is broadly recognized by independent experts as AGI (able to do essentially all the cognitive work of a top AI researcher/engineer); a developer's own claim does not count on its own. - Resolve by: 2027-02-19 - Verdict: Open - URL: https://sipoly.si/claim/kokotajlo-50pct-agi-by-2027/ ### At least 50%: AI builds a payment site from scratch, writes a hit-indistinguishable song and fine-tunes an LLM by itself by 2028 - Claimant: AI researchers (2023 expert survey) - Said: Jan 5, 2024 - Quote: "The aggregate forecasts give at least a 50% chance of AI systems achieving several milestones by 2028, including autonomously constructing a payment processing site from scratch, creating a song indistinguishable from a new song by a popular musician, and autonomously downloading and fine-tuning a large language model." - Source: https://arxiv.org/abs/2401.02843 - Resolution criteria (our interpretation): Our interpretation: YES if by end-2028 all three milestones are credibly demonstrated: an AI autonomously builds a working payment-processing site from scratch, generates a song judged indistinguishable from a new release by a popular artist, and autonomously downloads and fine-tunes an LLM. - Resolve by: 2028-12-31 - Verdict: Open - URL: https://sipoly.si/claim/ai-survey-2023-milestones-2028/ ### 10% chance all human occupations are fully automatable by 2037 (50% by 2116) - Claimant: AI researchers (2023 expert survey) - Said: Jan 5, 2024 - Quote: "However, the chance of all human occupations becoming fully automatable was forecast to reach 10% by 2037, and 50% as late as 2116" - Source: https://arxiv.org/abs/2401.02843 - Resolution criteria (our interpretation): Our interpretation: YES if by end-2037 every human occupation could be fully automated by AI at lower cost than human workers. - Resolve by: 2037-12-31 - Verdict: Open - URL: https://sipoly.si/claim/ai-survey-2023-full-automation-10pct-2037/ ### 50% chance machines outperform humans at every task by 2047 - Claimant: AI researchers (2023 expert survey) - Said: Jan 5, 2024 - Quote: "If science continues undisrupted, the chance of unaided machines outperforming humans in every possible task was estimated at 10% by 2027, and 50% by 2047." - Source: https://arxiv.org/abs/2401.02843 - Resolution criteria (our interpretation): Our interpretation: YES if by end-2047 unaided AI systems can do every task better and more cheaply than human workers. - Resolve by: 2047-12-31 - Verdict: Open - URL: https://sipoly.si/claim/ai-survey-2023-all-tasks-50pct-2047/ ### 10% chance machines outperform humans at every task by 2027 - Claimant: AI researchers (2023 expert survey) - Said: Jan 5, 2024 - Quote: "If science continues undisrupted, the chance of unaided machines outperforming humans in every possible task was estimated at 10% by 2027, and 50% by 2047." - Source: https://arxiv.org/abs/2401.02843 - Resolution criteria (our interpretation): Our interpretation: YES if by end-2027 unaided AI systems can do every task better and more cheaply than human workers (the survey's HLMI definition). - Resolve by: 2027-12-31 - Verdict: Open - URL: https://sipoly.si/claim/ai-survey-2023-all-tasks-10pct-2027/ ### 80%: Waymo opens public driverless rides in a new city in 2024 - Claimant: Vox Future Perfect - Said: Jan 1, 2024 - Quote: "Waymo will expand to a new city (80 percent)" - Source: https://www.vox.com/future-perfect/2024/1/1/24011179/2024-predictions-trump-politics-ohtani-oppenheimer-elections - Resolution criteria (our interpretation): Our interpretation: during 2024, ordinary members of the public can order driverless Waymo rides in at least one city beyond San Francisco and Phoenix. - Resolve by: 2024-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/vox-2024-waymo-new-city/ ### 75%: OpenAI releases "ChatGPT-5" by end of November 2024 - Claimant: Vox Future Perfect - Said: Jan 1, 2024 - Quote: "OpenAI will release ChatGPT-5 by the end of November 2024 (75 percent)" - Source: https://www.vox.com/future-perfect/2024/1/1/24011179/2024-predictions-trump-politics-ohtani-oppenheimer-elections - Resolution criteria (our interpretation): Our interpretation: OpenAI releases a product named ChatGPT-5 (or GPT-5) by 30 Nov 2024. - Resolve by: 2024-11-30 - Verdict: Wrong - URL: https://sipoly.si/claim/vox-2024-chatgpt5-by-november/ ### Gemini Ultra to developers and enterprises "early next year" - Claimant: Google / Google DeepMind - Said: Dec 6, 2023 - Quote: "we’ll make Gemini Ultra available to select customers, developers, partners and safety and responsibility experts for early experimentation and feedback before rolling it out to developers and enterprise customers early next year." - Source: https://blog.google/technology/ai/google-gemini-ai/ - Resolution criteria (our interpretation): Gemini Ultra broadly available by 30 Apr 2024. - Resolve by: 2024-04-30 - Verdict: Kept - URL: https://sipoly.si/claim/google-gemini-ultra-early-2024/ ### Within five years you will "simply tell your device" what you want instead of using apps - Claimant: Bill Gates - Said: Nov 9, 2023 - Quote: "In the next five years, this will change completely. You won’t have to use different apps for different tasks. You’ll simply tell your device, in everyday language, what you want to do." - Source: https://www.gatesnotes.com/AI-agents - Resolution criteria (our interpretation): Our interpretation: by 9 Nov 2028, mainstream consumers mostly accomplish multi-app tasks by telling an AI agent in natural language, rather than opening separate apps. - Resolve by: 2028-11-09 - Verdict: Open - URL: https://sipoly.si/claim/gates-agents-five-years-2023/ ### EO 14110: NIST generative-AI companion to the AI Risk Management Framework within 270 days - Claimant: The White House (U.S. executive branch) - Said: Oct 30, 2023 - Quote: "Within 270 days of the date of this order, to help ensure the development of safe, secure, and trustworthy AI systems, the Secretary of Commerce, acting through the Director of the National Institute of Standards and Technology (NIST)" - Source: https://www.federalregister.gov/documents/2023/11/01/2023-24283/safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence - Resolution criteria (our interpretation): Our interpretation: NIST publishes a final generative-AI companion to the AI RMF by 26 Jul 2024. - Resolve by: 2024-07-26 - Verdict: Kept - URL: https://sipoly.si/claim/eo14110-nist-genai-profile-270d/ ### EO 14110: launch a National AI Research Resource pilot within 90 days - Claimant: The White House (U.S. executive branch) - Said: Oct 30, 2023 - Quote: "Within 90 days of the date of this order, in coordination with the heads of agencies that the Director of NSF deems appropriate, launch a pilot program implementing the National AI Research Resource (NAIRR)" - Source: https://www.federalregister.gov/documents/2023/11/01/2023-24283/safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence - Resolution criteria (our interpretation): Our interpretation: the NSF launches a NAIRR pilot by 28 Jan 2024. - Resolve by: 2024-01-28 - Verdict: Kept - URL: https://sipoly.si/claim/eo14110-nairr-pilot-90d/ ### EO 14110: AI risk assessments for critical infrastructure within 90 days - Claimant: The White House (U.S. executive branch) - Said: Oct 30, 2023 - Quote: "Within 90 days of the date of this order, and at least annually thereafter, the head of each agency with relevant regulatory authority over critical infrastructure" - Source: https://www.federalregister.gov/documents/2023/11/01/2023-24283/safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence - Resolution criteria (our interpretation): Our interpretation: by 28 Jan 2024, the responsible agencies deliver AI risk assessments for critical-infrastructure sectors to DHS. - Resolve by: 2024-01-28 - Verdict: Kept - URL: https://sipoly.si/claim/eo14110-critical-infra-risk-90d/ ### EO 14110: frontier-model developers must report to the government within 90 days - Claimant: The White House (U.S. executive branch) - Said: Oct 30, 2023 - Quote: "Within 90 days of the date of this order, to ensure and verify the continuous availability of safe, reliable, and effective AI in accordance with the Defense Production Act" - Source: https://www.federalregister.gov/documents/2023/11/01/2023-24283/safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence - Resolution criteria (our interpretation): Our interpretation: by 28 Jan 2024, the Commerce Department is requiring developers of the most powerful models to report safety-test information under the Defense Production Act. - Resolve by: 2024-01-28 - Verdict: Kept - URL: https://sipoly.si/claim/eo14110-dpa-reporting-90d/ ### Global AI investment could approach $200 billion by 2025 - Claimant: Goldman Sachs Research - Said: Aug 2023 - Quote: "Goldman Sachs Research estimates AI investment could approach $100 billion in the U.S. and $200 billion globally by 2025." - Source: https://www.goldmansachs.com/insights/articles/ai-investment-forecast-to-approach-200-billion-globally-by-2025 - Resolution criteria (our interpretation): Our interpretation: correct if AI-related investment in 2025 was at least about $200B globally. - Resolve by: 2025-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/goldman-ai-investment-200b-2025/ ### An AI that turns $100k into $1M on its own "could be as little as two years away" - Claimant: Mustafa Suleyman - Said: Jul 14, 2023 - Quote: "Something like this could be as little as two years away." - Source: https://www.technologyreview.com/2023/07/14/1076296/mustafa-suleyman-my-new-turing-test-would-see-if-ai-can-make-1-million/ - Resolution criteria (our interpretation): Our interpretation: correct if by Jul 2025 an AI system is credibly documented as autonomously turning $100k into $1M on a retail platform within a few months. - Resolve by: 2025-07-14 - Verdict: Wrong - URL: https://sipoly.si/claim/suleyman-modern-turing-test-2y/ ### Solve the core technical challenges of superintelligence alignment in four years - Claimant: OpenAI - Said: Jul 5, 2023 - Quote: "Our goal is to solve the core technical challenges of superintelligence alignment in four years." - Source: https://openai.com/index/introducing-superalignment/ - Resolution criteria (our interpretation): Our interpretation: by July 2027, OpenAI's superalignment effort reports solving the core technical challenges of aligning superintelligence. - Resolve by: 2027-07-05 - Verdict: Withdrawn - URL: https://sipoly.si/claim/openai-superalignment-solved-4y/ ### 20% of secured compute to superalignment over four years - Claimant: OpenAI - Said: Jul 5, 2023 - Quote: "We are dedicating 20% of the compute we’ve secured to date over the next four years to solving the problem of superintelligence alignment." - Source: https://openai.com/index/introducing-superalignment/ - Resolution criteria (our interpretation): Our interpretation: the Superalignment effort receives 20% of the compute OpenAI had secured as of July 2023, over the period to July 2027. - Resolve by: 2027-07-05 - Verdict: Broken - URL: https://sipoly.si/claim/openai-superalignment-20pct/ ### "There will be no programmers in five years" - Claimant: Emad Mostaque - Said: Jul 3, 2023 - Quote: "There will be no programmers in five years." - Source: https://decrypt.co/147191/no-human-programmers-five-years-ai-stability-ceo - Resolution criteria (our interpretation): Our interpretation: by Jul 2028, professional software developer employment has largely disappeared (e.g. down 90%+). - Resolve by: 2028-07-31 - Verdict: Open - URL: https://sipoly.si/claim/mostaque-no-programmers/ ### By 2030 AI will outcompete most professional mathematicians at research - Claimant: Jacob Steinhardt - Said: Jun 7, 2023 - Quote: "In a domain like mathematics research where work can be checked automatically, I’d predict that GPT 2030 will outcompete most professional mathematicians." - Source: https://bounded-regret.ghost.io/what-will-gpt-2030-look-like/ - Resolution criteria (our interpretation): Our interpretation: by end-2030, AI systems routinely produce publishable mathematical research at or above the level of a typical professional mathematician, as judged by mainstream mathematicians. - Resolve by: 2030-12-31 - Verdict: Open - URL: https://sipoly.si/claim/steinhardt-gpt2030-outcompete-mathematicians/ ### "GPT 2030" will likely be superhuman at coding, hacking and math - Claimant: Jacob Steinhardt - Said: Jun 7, 2023 - Quote: "GPT 2030 will likely be superhuman at various specific tasks, including coding, hacking, and math, and potentially protein design" - Source: https://bounded-regret.ghost.io/what-will-gpt-2030-look-like/ - Resolution criteria (our interpretation): Our interpretation: by end-2030, the best available AI system beats top human experts on standard measures of coding (e.g. top competitive-programming ranks), offensive security, and mathematics. - Resolve by: 2030-12-31 - Verdict: Open - URL: https://sipoly.si/claim/steinhardt-gpt2030-superhuman-coding-math/ ### Transformative AGI by 2043 is less than 1% likely - Claimant: Ted Sanders - Said: Jun 6, 2023 - Quote: "we estimate the likelihood of transformative artificial general intelligence (AGI) by 2043 and find it to be <1%." - Source: https://www.lesswrong.com/posts/DgzdLzDGsqoRXhCK7/transformative-agi-by-2043-is-less-than-1-likely - Resolution criteria (our interpretation): Our interpretation: YES if by end-2043 AI can perform nearly all economically valuable tasks at or below human cost (the contest definition). - Resolve by: 2043-12-31 - Verdict: Open - URL: https://sipoly.si/claim/sanders-transformative-agi-2043-under-1pct/ ### 30% of IBM back-office roles replaced over five years - Claimant: Arvind Krishna - Said: May 1, 2023 - Quote: "I could easily see 30% of that getting replaced by AI and automation over a five-year period." - Source: https://fortune.com/2023/05/01/ibm-ceo-ai-artificial-intelligence-back-office-jobs-pause-hiring/ - Resolution criteria (our interpretation): By May 2028, ~30% (~7,800) of those IBM roles replaced by AI/automation, per IBM disclosures or credible reporting. - Resolve by: 2028-05-01 - Verdict: Open - URL: https://sipoly.si/claim/krishna-30pct-back-office/ ### Generative AI could raise global GDP by 7% over ten years - Claimant: Goldman Sachs Research - Said: Apr 5, 2023 - Quote: "they could drive a 7% (or almost $7 trillion) increase in global GDP and lift productivity growth by 1.5 percentage points over a 10-year period." - Source: https://www.goldmansachs.com/insights/articles/generative-ai-could-raise-global-gdp-by-7-percent - Resolution criteria (our interpretation): Our interpretation: by 2033, US/global productivity growth runs about 1.5 percentage points above its pre-2023 trend, as estimated by mainstream sources. Hedged ("could"). - Resolve by: 2033-12-31 - Verdict: Open - URL: https://sipoly.si/claim/goldman-genai-gdp-7pct-10y/ ### Bet: AI will not get an A on 5 of his 6 latest midterms by January 2029 - Claimant: Bryan Caplan - Said: Jan 19, 2023 - Quote: "If the AI gets an A on at least 5 of out 6 of those exams using same grading scale as his students, then Bryan owes Matthew $500." - Source: https://www.betonit.ai/p/ai-bet - Resolution criteria (our interpretation): Our interpretation: Caplan is right if an AI does not get an A (A- counts) on at least 5 of his 6 most recent midterms by 30 Jan 2029, under the bet's terms. - Resolve by: 2029-01-30 - Verdict: Open - URL: https://sipoly.si/claim/caplan-ai-no-a-midterms-2029/ ### 45% that Robin Hanson wins his bet that GPT revenue stays under $1B - Claimant: XPT domain experts - Said: Oct 2022 - Quote: "35. GPT Revenue (Hanson Wins Bet that GPT Revenue < $1B) 45% 0% 0.90 6" - Source: https://forecastingresearch.org/research/near-term-xpt-accuracy - Resolution criteria (our interpretation): Our interpretation: YES if Hanson wins his bet that GPT revenue stays below $1B (per the XPT question). - Resolve by: 2025-06-30 - Verdict: Correct - URL: https://sipoly.si/claim/xpt-experts-gpt-revenue-hanson/ ### 53.5% that Robin Hanson wins his bet that GPT revenue stays under $1B - Claimant: XPT superforecasters - Said: Oct 2022 - Quote: "35. GPT Revenue (Hanson Wins Bet that GPT Revenue < $1B) 53.5% 0% 1.07 32" - Source: https://forecastingresearch.org/research/near-term-xpt-accuracy - Resolution criteria (our interpretation): Our interpretation: YES if Hanson wins his bet that GPT revenue stays below $1B (per the XPT question). - Resolve by: 2025-06-30 - Verdict: Wrong - URL: https://sipoly.si/claim/xpt-supers-gpt-revenue-hanson/ ### Median forecast: largest ML model would have 150 trillion parameters by 2024 (actual ~10 trillion) - Claimant: XPT domain experts - Said: Oct 2022 - Quote: "49. Largest Number of Parameters in a Machine Learning Model 150 trillion 10 trillion 2.74 7" - Source: https://forecastingresearch.org/research/near-term-xpt-accuracy - Resolution criteria (our interpretation): Our interpretation: treated as a point forecast; correct if within a factor of 2 of the resolved value. - Resolve by: 2024-12-31 - Verdict: Wrong - URL: https://sipoly.si/claim/xpt-experts-largest-params-2024/ ### Median forecast: largest ML model would have 100 trillion parameters by 2024 (actual ~10 trillion) - Claimant: XPT superforecasters - Said: Oct 2022 - Quote: "49. Largest Number of Parameters in a Machine Learning Model 100 trillion 10 trillion 1.71 31" - Source: https://forecastingresearch.org/research/near-term-xpt-accuracy - Resolution criteria (our interpretation): Our interpretation: treated as a point forecast; correct if within a factor of 2 of the resolved value. - Resolve by: 2024-12-31 - Verdict: Wrong - URL: https://sipoly.si/claim/xpt-supers-largest-params-2024/ ### Median forecast for the maximum compute used in an AI experiment by 2024 was about 6x too low - Claimant: XPT superforecasters - Said: Oct 2022 - Quote: "45. Maximum Compute Used in an AI Experiment 100,000 578,703.7 1.92 33" - Source: https://forecastingresearch.org/research/near-term-xpt-accuracy - Resolution criteria (our interpretation): Our interpretation: treated as a point forecast; correct if within a factor of 2 of the resolved value. - Resolve by: 2024-12-31 - Verdict: Wrong - URL: https://sipoly.si/claim/xpt-supers-max-compute-2024/ ### Median forecast: MATH benchmark state of the art 71% by end of 2024 (actual 87.9%) - Claimant: XPT superforecasters - Said: Oct 2022 - Quote: "39. MATH Dataset Benchmark 71% 87.92% 1.38 30" - Source: https://forecastingresearch.org/research/near-term-xpt-accuracy - Resolution criteria (our interpretation): Our interpretation: treated as a point forecast of the best MATH score by end-2024; correct if within 5 points of the resolved value. - Resolve by: 2024-12-31 - Verdict: Wrong - URL: https://sipoly.si/claim/xpt-supers-math-2024/ ### Median forecast: MMLU state of the art 77.75% by end of 2024 (actual 88.7%) - Claimant: XPT superforecasters - Said: Oct 2022 - Quote: "40. “Massive Multitask Language Understanding” Benchmark 77.75% 88.7% 1.59 32" - Source: https://forecastingresearch.org/research/near-term-xpt-accuracy - Resolution criteria (our interpretation): Our interpretation: treated as a point forecast of the best MMLU score by end-2024; correct if within 5 points of the resolved value. - Resolve by: 2024-12-31 - Verdict: Wrong - URL: https://sipoly.si/claim/xpt-supers-mmlu-2024/ ### 8.6%: AI gold-level performance at the IMO by 2025 - Claimant: XPT domain experts - Said: Oct 2022 - Quote: "an outcome to which domain experts assigned only an 8.6% probability and superforecasters a mere 2.3% probability." - Source: https://forecastingresearch.org/research/near-term-xpt-accuracy - Resolution criteria (our interpretation): Our interpretation: YES if an AI system achieves gold-medal-level performance at the International Mathematical Olympiad by 2025. - Resolve by: 2025-12-31 - Verdict: Wrong - URL: https://sipoly.si/claim/xpt-experts-imo-gold-2025/ ### 2.3%: AI gold-level performance at the IMO by 2025 - Claimant: XPT superforecasters - Said: Oct 2022 - Quote: "an outcome to which domain experts assigned only an 8.6% probability and superforecasters a mere 2.3% probability." - Source: https://forecastingresearch.org/research/near-term-xpt-accuracy - Resolution criteria (our interpretation): Our interpretation: YES if an AI system achieves gold-medal-level performance at the International Mathematical Olympiad by 2025. - Resolve by: 2025-12-31 - Verdict: Wrong - URL: https://sipoly.si/claim/xpt-supers-imo-gold-2025/ ### Median ~2040 for transformative AI (50% by 2040) - Claimant: Ajeya Cotra - Said: Aug 2, 2022 - Quote: "A median of ~2040 (a decrease of ~10 years from 2050)." - Source: https://www.lesswrong.com/posts/AfH2oPHCApdKicM4m/two-year-update-on-my-personal-ai-timelines - Resolution criteria (our interpretation): Our interpretation: YES if transformative AI (as defined in her report) exists by 31 Dec 2040. - Resolve by: 2040-12-31 - Verdict: Open - URL: https://sipoly.si/claim/cotra-tai-median-2040/ ### ~35% probability of transformative AI by 2036 - Claimant: Ajeya Cotra - Said: Aug 2, 2022 - Quote: "~35% probability by 2036 (a ~3x likelihood ratio [3] vs 15%)." - Source: https://www.lesswrong.com/posts/AfH2oPHCApdKicM4m/two-year-update-on-my-personal-ai-timelines - Resolution criteria (our interpretation): Our interpretation: YES if transformative AI (as defined in her report) exists by 31 Dec 2036. - Resolve by: 2036-12-31 - Verdict: Open - URL: https://sipoly.si/claim/cotra-tai-2036-35pct/ ### ~15% probability of transformative AI by 2030 - Claimant: Ajeya Cotra - Said: Aug 2, 2022 - Quote: "~15% probability by 2030 (a decrease of ~6 years from 2036)." - Source: https://www.lesswrong.com/posts/AfH2oPHCApdKicM4m/two-year-update-on-my-personal-ai-timelines - Resolution criteria (our interpretation): Our interpretation: YES if transformative AI (as defined in her report) exists by 31 Dec 2030. - Resolve by: 2030-12-31 - Verdict: Open - URL: https://sipoly.si/claim/cotra-tai-2030-15pct/ ### >16% that AI has IMO-gold capability by end of 2025 - Claimant: Eliezer Yudkowsky - Said: Feb 26, 2022 - Quote: "I'll stand by a >16% probability of the technical capability existing by end of 2025" - Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challenge-bet-with-eliezer - Resolution criteria (our interpretation): Technical capability for IMO gold (grand-challenge conditions) exists by end of 2025. - Resolve by: 2025-12-31 - Verdict: Wrong - URL: https://sipoly.si/claim/yudkowsky-imo-16pct/ ### 4% that AI solves the hardest IMO problem by 2025 - Claimant: Paul Christiano - Said: Feb 26, 2022 - Quote: "I'd put 4% on "For the 2022, 2023, 2024, or 2025 IMO an AI built before the IMO is able to solve the single hardest problem"" - Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challenge-bet-with-eliezer - Resolution criteria (our interpretation): An AI solves the designated hardest problem (usually #6) at a 2022-2025 IMO. - Resolve by: 2025-07-31 - Verdict: Correct - URL: https://sipoly.si/claim/christiano-imo-hardest-4pct/ ### 8% that AI gets IMO gold by 2025 - Claimant: Paul Christiano - Said: Feb 26, 2022 - Quote: "Maybe I'll go 8% on "gets gold" instead of "solves hardest problem."" - Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challenge-bet-with-eliezer - Resolution criteria (our interpretation): An AI built before the IMO achieves a gold-medal score at the 2022-2025 IMO. - Resolve by: 2025-07-31 - Verdict: Wrong - URL: https://sipoly.si/claim/christiano-imo-gold-8pct/ ### Apollo Go robotaxis in 65 cities by 2025 (100 by 2030) - Claimant: Baidu - Said: Nov 17, 2021 - Quote: "Apollo Go, Baidu’s robotaxi service, aims to be in 65 cities by 2025 and 100 cities by 2030, the firm’s co-founder and CEO Robin Li said on an analyst call Wednesday." - Source: https://techcrunch.com/2021/11/17/baidu-robotaxi-2030/ - Resolution criteria (our interpretation): Our interpretation: Apollo Go operates (testing or commercial) in at least 65 cities by 31 Dec 2025. - Resolve by: 2025-12-31 - Verdict: Broken - URL: https://sipoly.si/claim/baidu-apollo-65-cities-2025/ ### Fully driverless public robotaxi service with Lyft in Las Vegas in 2023 - Claimant: Motional - Said: Nov 9, 2021 - Quote: "Motional’s next-generation robotaxis, the all-electric Hyundai IONIQ 5-based robotaxi, will be available on the Lyft app in Las Vegas, starting in 2023." - Source: https://www.lyft.com/blog/posts/motional-and-lyft-to-launch-fully-driverless-ride-hail-service-in-las-vegas - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2023, Lyft riders in Las Vegas can hail Motional robotaxis with no safety operator on board. - Resolve by: 2023-12-31 - Verdict: Broken - URL: https://sipoly.si/claim/motional-lyft-driverless-vegas-2023/ ### Cruise robotaxis: potential for $50 billion in annual revenue by end of decade - Claimant: General Motors - Said: Oct 6, 2021 - Quote: "With Cruise, GM has a market-leading position in autonomous services with the potential to deliver $50 billion in revenue annually by the end of the decade." - Source: https://investor.gm.com/news-releases/news-release-details/gm-details-plan-double-its-revenue-drive-even-higher-margins - Resolution criteria (our interpretation): Our interpretation: Cruise's autonomous services reach about $50B in annual revenue by 2030. - Resolve by: 2030-12-31 - Verdict: Withdrawn - URL: https://sipoly.si/claim/gm-cruise-50b-revenue-2030/ ### A new chip shortage in 2026 - Claimant: Daniel Kokotajlo - Said: Aug 6, 2021 - Quote: "We’re in a new chip shortage. Just when the fabs thought they had caught up to demand…" - Source: https://www.lesswrong.com/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-like - Resolution criteria (our interpretation): Our interpretation: during 2026, widely reported AI-driven shortages of leading-edge chips or memory. - Resolve by: 2026-12-31 - Verdict: Open - URL: https://sipoly.si/claim/kokotajlo-chip-shortage-2026/ ### By 2025, AIs play Diplomacy as well as human experts - Claimant: Daniel Kokotajlo - Said: Aug 6, 2021 - Quote: "After years of tinkering and incremental progress, AIs can now play Diplomacy as well as human experts ." - Source: https://www.lesswrong.com/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-like - Resolution criteria (our interpretation): By end of 2025, an AI plays (no-press or full-press) Diplomacy at human-expert level. - Resolve by: 2025-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/kokotajlo-diplomacy-2025/ ### Long Bet: AI will not lift US productivity growth above 1.8%/yr in the 2020s - Claimant: Robert J. Gordon - Said: 2021 - Quote: "Will robots and AI bring a new revival comparable to the post-1995 digital revival? A reason for doubt is that a doubling of the U.S. stock of robots in the past decade failed to revive manufacturing productivity growth" - Source: https://longbets.org/868/ - Resolution criteria (our interpretation): Our interpretation: Gordon is right if BLS private nonfarm business productivity grows at an average of 1.8%/yr or less from 2020Q1 to 2029Q4. - Resolve by: 2030-03-31 - Verdict: Open - URL: https://sipoly.si/claim/gordon-no-productivity-revival-2029/ ### Long Bet: US productivity growth will average over 1.8%/yr in 2020-2029 - Claimant: Erik Brynjolfsson - Said: 2021 - Quote: "Private Nonfarm business productivity growth will average over 1.8 percent per year from the first quarter (Q1) of 2020 to the last quarter of 2029 (Q4)." - Source: https://longbets.org/868/ - Resolution criteria (our interpretation): Our interpretation: as in the bet terms, BLS private nonfarm business labor productivity grows at an average above 1.8%/yr from 2020Q1 to 2029Q4. - Resolve by: 2030-03-31 - Verdict: Open - URL: https://sipoly.si/claim/brynjolfsson-productivity-1-8pct-2029/ ### Basic functionality for Level 5 autonomy complete in 2020 - Claimant: Elon Musk - Said: Jul 9, 2020 - Quote: "I remain confident that we will have the basic functionality for level five autonomy complete this year." - Source: https://www.bbc.com/news/technology-53349313 - Resolution criteria (our interpretation): Our interpretation: Tesla ships or demonstrates no-supervision (Level 5-class) driving functionality by 31 Dec 2020. - Resolve by: 2020-12-31 - Verdict: Wrong - URL: https://sipoly.si/claim/musk-level5-2020/ ### "Over a million robotaxis on the road" in 2020 - Claimant: Elon Musk - Said: Apr 22, 2019 - Quote: "From our standpoint, if you fast forward a year, maybe a year and three months, but next year for sure, we’ll have over a million robotaxis on the road" - Source: https://techcrunch.com/2019/04/22/tesla-plans-to-launch-a-robotaxi-network-in-2020/ - Resolution criteria (our interpretation): Over one million Tesla vehicles operating as driverless robotaxis by 31 Dec 2020 (generous: by Jul 2021). - Resolve by: 2021-07-31 - Verdict: Wrong - URL: https://sipoly.si/claim/musk-million-robotaxis-2020/ ### AI will displace about 40% of the world's jobs within 15 years - Claimant: Kai-Fu Lee - Said: Jan 13, 2019 - Quote: "all together in 15 years, that's going to displace about 40% of the jobs in the world." - Source: https://www.cbsnews.com/news/60-minutes-ai-facial-and-emotional-recognition-how-kai-fu-lee-is-advancing-artificial-intelligence-2019-07-14/ - Resolution criteria (our interpretation): Our interpretation: by Jan 2034, roughly 40% of jobs worldwide have been displaced by AI and automation (eliminated or replaced, not just changed). - Resolve by: 2034-01-13 - Verdict: Open - URL: https://sipoly.si/claim/kai-fu-lee-40pct-jobs-15y/ ### Up to 20,000 Jaguar I-PACEs "in the next few years" - Claimant: Waymo - Said: Mar 27, 2018 - Quote: "We’ll add up to 20,000 I-PACEs to Waymo’s fleet in the next few years" - Source: https://waymo.com/blog/2018/03/meet-our-newest-self-driving-vehicle - Resolution criteria (our interpretation): Our interpretation: fleet additions approaching 20,000 I-PACEs within five years (by Mar 2023). - Resolve by: 2023-03-27 - Verdict: Broken - URL: https://sipoly.si/claim/waymo-20000-ipace/ ### 50%: AI risk will feel more widely accepted as a field by 2023 (40% same, 10% less) - Claimant: Scott Alexander - Said: Feb 2018 - Quote: "7. AI risk as a field subjectively feels more/same/less widely accepted than today: 50%/40%/10%" - Source: https://slatestarcodex.com/2018/02/15/five-more-years/ - Resolution criteria (our interpretation): Our interpretation: YES if by 2023 AI risk is more widely accepted as a field than in 2018. - Resolve by: 2023-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/ssc-2018-ai-risk-more-accepted/ ### 80%: MIRI still exists in 2023 - Claimant: Scott Alexander - Said: Feb 2018 - Quote: "6. MIRI still exists in 2023: 80%" - Source: https://slatestarcodex.com/2018/02/15/five-more-years/ - Resolution criteria (our interpretation): Our interpretation: YES if the Machine Intelligence Research Institute still exists as an organization in 2023. - Resolve by: 2023-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/ssc-2018-miri-exists-2023/ ### 70%: AI beats a top human player at StarCraft by 2023 - Claimant: Scott Alexander - Said: Feb 15, 2018 - Quote: "AI beats a top human player at Starcraft: 70%" - Source: https://slatestarcodex.com/2018/02/15/five-more-years/ - Resolution criteria (our interpretation): By 15 Feb 2023 an AI defeats a top professional StarCraft player. - Resolve by: 2023-02-15 - Verdict: Correct - URL: https://sipoly.si/claim/ssc-2018-starcraft/ ### 30%: average person can buy a self-driving car for under $100,000 by 2023 - Claimant: Scott Alexander - Said: Feb 15, 2018 - Quote: "Average person can buy a self-driving car for less than $100,000: 30%" - Source: https://slatestarcodex.com/2018/02/15/five-more-years/ - Resolution criteria (our interpretation): By 15 Feb 2023 a consumer car capable of unsupervised driving is purchasable under $100k. - Resolve by: 2023-02-15 - Verdict: Correct - URL: https://sipoly.si/claim/ssc-2018-car-under-100k/ ### 10%: 5% of US truck drivers replaced by self-driving trucks by 2023 - Claimant: Scott Alexander - Said: Feb 15, 2018 - Quote: "At least 5% of US truck drivers have been replaced by self-driving trucks: 10%" - Source: https://slatestarcodex.com/2018/02/15/five-more-years/ - Resolution criteria (our interpretation): By 15 Feb 2023, ≥5% of US truck drivers replaced by autonomous trucks. - Resolve by: 2023-02-15 - Verdict: Correct - URL: https://sipoly.si/claim/ssc-2018-trucks-5pct/ ### 30%: self-driving hail in five of the ten largest US cities by 2023 - Claimant: Scott Alexander - Said: Feb 15, 2018 - Quote: "…in at least five of ten largest US cities: 30%" - Source: https://slatestarcodex.com/2018/02/15/five-more-years/ - Resolution criteria (our interpretation): By 15 Feb 2023 public driverless hail available in 5 of the 10 largest US cities. - Resolve by: 2023-02-15 - Verdict: Correct - URL: https://sipoly.si/claim/ssc-2018-selfdriving-five-cities/ ### 80%: average person can hail a self-driving car in at least one US city by 2023 - Claimant: Scott Alexander - Said: Feb 15, 2018 - Quote: "Average person can hail a self-driving car in at least one US city: 80%" - Source: https://slatestarcodex.com/2018/02/15/five-more-years/ - Resolution criteria (our interpretation): By 15 Feb 2023 a member of the public can hail a driverless car in a US city. - Resolve by: 2023-02-15 - Verdict: Correct - URL: https://sipoly.si/claim/ssc-2018-selfdriving-one-city/ ### A robot as intelligent, attentive and faithful as a dog: not before 2048 - Claimant: Rodney Brooks - Said: Jan 1, 2018 - Quote: "A robot that seems as intelligent, as attentive, and as faithful, as a dog. NET 2048" - Source: https://rodneybrooks.com/my-dated-predictions/ - Resolution criteria (our interpretation): Our interpretation: correct if no robot is broadly judged as intelligent, attentive and faithful as a dog before 1 Jan 2048. - Resolve by: 2047-12-31 - Verdict: Open - URL: https://sipoly.si/claim/brooks-dog-level-robot-net2048/ ### AI with an ongoing existence at the level of a mouse: not before 2030 - Claimant: Rodney Brooks - Said: Jan 1, 2018 - Quote: "An AI system with an ongoing existence (no day is the repeat of another day as it currently is for all AI systems) at the level of a mouse. NET 2030" - Source: https://rodneybrooks.com/my-dated-predictions/ - Resolution criteria (our interpretation): Our interpretation: correct if before 1 Jan 2030 no AI system is generally accepted as having a continuous, learning, mouse-level existence. - Resolve by: 2029-12-31 - Verdict: Open - URL: https://sipoly.si/claim/brooks-mouse-level-ai-net2030/ ### Driverless taxi with arbitrary pick-up and drop-off in a major US city: not before 2032 - Claimant: Rodney Brooks - Said: Jan 1, 2018 - Quote: "A driverless "taxi" service in a major US city with arbitrary pick and drop off locations, even in a restricted geographical area. NET 2032" - Source: https://rodneybrooks.com/my-dated-predictions/ - Resolution criteria (our interpretation): Our interpretation: correct if no driverless taxi service with arbitrary pick-up/drop-off points operates in a major US city before 1 Jan 2032; wrong if one does. - Resolve by: 2031-12-31 - Verdict: Wrong - URL: https://sipoly.si/claim/brooks-arbitrary-pickup-taxi-net2032/ ### Deployed conversational agent with long-term context by 2025 - Claimant: Rodney Brooks - Said: Jan 1, 2018 - Quote: "A conversational agent that both carries long term context, and does not easily fall into recognizable and repeated patterns. Lab demo: NET 2023 Deployed systems: 2025" - Source: https://rodneybrooks.com/my-dated-predictions/ - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2025, a widely deployed conversational agent carries long-term context across conversations and does not fall into recognisable repeated patterns. - Resolve by: 2025-12-31 - Verdict: Wrong - URL: https://sipoly.si/claim/brooks-conversational-agent-2025/ ### Lab demo of a robot doing the "last 10 yards" of package delivery: not before 2025 - Claimant: Rodney Brooks - Said: Jan 1, 2018 - Quote: "A robot that can carry out the last 10 yards of delivery, getting from a vehicle into a house and putting the package inside the front door. Lab demo: NET 2025" - Source: https://rodneybrooks.com/my-dated-predictions/ - Resolution criteria (our interpretation): Our interpretation: correct if no robot was demonstrated, even in a lab, carrying a package from a vehicle into a house before 1 Jan 2025. - Resolve by: 2024-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/brooks-last-10-yards-demo-net2025/ ### Lab demo of a robot that can navigate almost any US home: not before 2026 - Claimant: Rodney Brooks - Said: Jan 1, 2018 - Quote: "A robot that can navigate around just about any US home, with its steps, its clutter, its narrow pathways between furniture, etc. Lab demo: NET 2026" - Source: https://rodneybrooks.com/my-dated-predictions/ - Resolution criteria (our interpretation): Our interpretation: correct if no lab demonstration of a home-class robot navigating a cluttered home including steps took place before 1 Jan 2026. - Resolve by: 2025-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/brooks-home-robot-lab-demo-net2026/ ### A major city bans human-driven cars from part of town for driverless cars: 2027 to 2031 - Claimant: Rodney Brooks - Said: Jan 1, 2018 - Quote: "A major city bans parking and cars with drivers from a non-trivial portion of a city so that driverless cars have free reign in that area. NET 2027 BY 2031" - Source: https://rodneybrooks.com/my-dated-predictions/ - Resolution criteria (our interpretation): Our interpretation: correct if a major city first bans parking and human-driven cars from a non-trivial area in favour of driverless cars in 2027-2031 (not before 2027, and by end of 2031). - Resolve by: 2031-12-31 - Verdict: Open - URL: https://sipoly.si/claim/brooks-city-bans-human-drivers-net2027/ ### First freeway lane reserved for truly driverless cars: not before 2021 - Claimant: Rodney Brooks - Said: Jan 1, 2018 - Quote: "First dedicated lane where only cars in truly driverless mode are allowed on a public freeway. NET 2021" - Source: https://rodneybrooks.com/my-dated-predictions/ - Resolution criteria (our interpretation): Our interpretation: correct if no public freeway lane reserved for driverless-mode cars opened before 1 Jan 2021. - Resolve by: 2020-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/brooks-driverless-freeway-lane/ ### Dual-use (driver/driverless) taxi service in 10 major US cities: not before 2025 - Claimant: Rodney Brooks - Said: Jan 1, 2018 - Quote: "Such "taxi" services where the cars are also used with drivers at other times and with extended geography, in 10 major US cities NET 2025" - Source: https://rodneybrooks.com/my-dated-predictions/ - Resolution criteria (our interpretation): Our interpretation: correct if before 1 Jan 2025 no driverless taxi service, using cars that are also driven by people at other times, operated in 10 major US cities. - Resolve by: 2024-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/brooks-dual-use-taxi-10-cities-net2025/ ### Deployed "last 10 yards" delivery robots: not before 2028 - Claimant: Rodney Brooks - Said: Jan 1, 2018 - Quote: "Lab demo: NET 2025 Deployed systems: NET 2028" - Source: https://rodneybrooks.com/my-dated-predictions/ - Resolution criteria (our interpretation): Our interpretation: correct if no commercially deployed robot does vehicle-to-front-door package delivery before 1 Jan 2028. - Resolve by: 2027-12-31 - Verdict: Open - URL: https://sipoly.si/claim/brooks-last-10-yards-deployed-net2028/ ### Dexterous robot hands generally available: not before 2030 (hopefully by 2040) - Claimant: Rodney Brooks - Said: Jan 1, 2018 - Quote: "Dexterous robot hands generally available. NET 2030 BY 2040 (I hope!)" - Source: https://rodneybrooks.com/my-dated-predictions/ - Resolution criteria (our interpretation): Our interpretation: correct if articulated, human-like dexterous robot hands are not generally available as deployed products before 1 Jan 2030. - Resolve by: 2029-12-31 - Verdict: Open - URL: https://sipoly.si/claim/brooks-dexterous-hands-net2030/ ### Affordable home-navigating robot: not before 2035 - Claimant: Rodney Brooks - Said: Jan 1, 2018 - Quote: "Expensive product: NET 2030 Affordable product: NET 2035 What is easy for humans is still very, very hard for robots." - Source: https://rodneybrooks.com/my-dated-predictions/ - Resolution criteria (our interpretation): Our interpretation: correct if no affordable (mass-market consumer price) robot able to navigate almost any US home is on sale before 1 Jan 2035. - Resolve by: 2034-12-31 - Verdict: Open - URL: https://sipoly.si/claim/brooks-home-robot-affordable-net2035/ ### Home-navigating robot as an expensive product: not before 2030 - Claimant: Rodney Brooks - Said: Jan 1, 2018 - Quote: "A robot that can navigate around just about any US home, with its steps, its clutter, its narrow pathways between furniture, etc. Lab demo: NET 2026 Expensive product: NET 2030" - Source: https://rodneybrooks.com/my-dated-predictions/ - Resolution criteria (our interpretation): Our interpretation: correct if no robot able to navigate almost any US home (steps, clutter, narrow paths) is sold as a product, at any price, before 1 Jan 2030. - Resolve by: 2029-12-31 - Verdict: Open - URL: https://sipoly.si/claim/brooks-home-robot-expensive-net2030/ ### Multi-task physical-assistance robot for the elderly: not before 2028 - Claimant: Rodney Brooks - Said: Jan 1, 2018 - Quote: "A robot that can provide physical assistance to the elderly over multiple tasks (e.g., getting into and out of bed, washing, using the toilet, etc.) rather than just a point solution. NET 2028" - Source: https://rodneybrooks.com/predictions-scorecard-2025-january-01/ - Resolution criteria (our interpretation): Correct if no such multi-task elder-assistance robot is available before 1 Jan 2028. - Resolve by: 2027-12-31 - Verdict: Open - URL: https://sipoly.si/claim/brooks-elder-care-robot/ ### Generally agreed "next big thing" beyond deep learning: by 2027 - Claimant: Rodney Brooks - Said: Jan 1, 2018 - Quote: "Emergence of the generally agreed upon "next big thing" in AI beyond deep learning. NET 2023" - Source: https://rodneybrooks.com/predictions-scorecard-2025-january-01/ - Resolution criteria (our interpretation): Our interpretation: by end of 2027, a widely agreed successor paradigm to deep learning has emerged. - Resolve by: 2027-12-31 - Verdict: Open - URL: https://sipoly.si/claim/brooks-next-big-thing/ ### Driverless taxi service in 50 of the 100 biggest US cities: not before 2028 - Claimant: Rodney Brooks - Said: Jan 1, 2018 - Quote: "Such "taxi" service as above in 50 of the 100 biggest US cities. NET 2028" - Source: https://rodneybrooks.com/predictions-scorecard-2025-january-01/ - Resolution criteria (our interpretation): Correct if no such service exists in 50 of the 100 largest US cities before 1 Jan 2028. - Resolve by: 2027-12-31 - Verdict: Open - URL: https://sipoly.si/claim/brooks-taxi-50-cities/ ### China to have initial AI laws, ethics norms and safety assessment by 2025 - Claimant: China State Council - Said: Jul 20, 2017 - Quote: "By 2025 China will have seen the initial establishment of AI laws and regulations, ethical norms and policy systems, and the formation of AI security assessment and control capabilities." - Source: https://digichina.stanford.edu/work/full-translation-chinas-new-generation-artificial-intelligence-development-plan-2017/ - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2025, China has binding AI-specific regulations in force plus official AI ethics norms and a security-assessment regime. - Resolve by: 2025-12-31 - Verdict: Kept - URL: https://sipoly.si/claim/china-aidp-laws-2025/ ### China's core AI industry to exceed 400 billion RMB by 2025 - Claimant: China State Council - Said: Jul 20, 2017 - Quote: "the scale of AI’s core industry will be more than 400 billion RMB, and the scale of related industries will exceed 5 trillion RMB." - Source: https://digichina.stanford.edu/work/full-translation-chinas-new-generation-artificial-intelligence-development-plan-2017/ - Resolution criteria (our interpretation): Our interpretation: official Chinese statistics (e.g. MIIT or CAICT) put the core AI industry above 400 billion RMB for 2025. - Resolve by: 2025-12-31 - Verdict: Kept - URL: https://sipoly.si/claim/china-aidp-core-industry-2025/ ### China to be "the world's primary AI innovation center" by 2030 - Claimant: China State Council - Said: Jul 20, 2017 - Quote: "Third, by 2030, China’s AI theories, technologies, and applications should achieve world-leading levels, making China the world’s primary AI innovation center" - Source: https://digichina.stanford.edu/work/full-translation-chinas-new-generation-artificial-intelligence-development-plan-2017/ - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2030, China leads the world in AI research and frontier capability (e.g. top frontier models and most highly-cited AI research), per major independent indices such as the Stanford AI Index. - Resolve by: 2030-12-31 - Verdict: Pending - URL: https://sipoly.si/claim/china-aidp-world-leader-2030/ ### AI will outperform humans at translating languages by 2024 (median) - Claimant: ML researchers (2016 expert survey) - Said: May 24, 2017 - Quote: "Researchers predict AI will outperform humans in many activities in the next ten years, such as translating languages (by 2024), writing high-school essays (by 2026), driving a truck (by 2027), working in retail (by 2031), writing a bestselling book (by 2049), and working as a surgeon (by 2053)." - Source: https://arxiv.org/abs/1705.08807 - Resolution criteria (our interpretation): Our interpretation: by end-2024, machine translation matches or beats fluent amateur human translators across major language pairs. - Resolve by: 2024-12-31 - Verdict: Correct - URL: https://sipoly.si/claim/ml-survey-2016-translate-2024/ ### AI will outperform humans at working as a surgeon by 2053 (median) - Claimant: ML researchers (2016 expert survey) - Said: May 24, 2017 - Quote: "Researchers predict AI will outperform humans in many activities in the next ten years, such as translating languages (by 2024), writing high-school essays (by 2026), driving a truck (by 2027), working in retail (by 2031), writing a bestselling book (by 2049), and working as a surgeon (by 2053)." - Source: https://arxiv.org/abs/1705.08807 - Resolution criteria (our interpretation): Our interpretation: by end-2053, autonomous surgical systems perform at least as well as human surgeons across common operations. - Resolve by: 2053-12-31 - Verdict: Open - URL: https://sipoly.si/claim/ml-survey-2016-surgeon-2053/ ### AI will write a bestselling book by 2049 (median) - Claimant: ML researchers (2016 expert survey) - Said: May 24, 2017 - Quote: "Researchers predict AI will outperform humans in many activities in the next ten years, such as translating languages (by 2024), writing high-school essays (by 2026), driving a truck (by 2027), working in retail (by 2031), writing a bestselling book (by 2049), and working as a surgeon (by 2053)." - Source: https://arxiv.org/abs/1705.08807 - Resolution criteria (our interpretation): Our interpretation: by end-2049, an AI-written book (credited as such) reaches the New York Times bestseller list. - Resolve by: 2049-12-31 - Verdict: Open - URL: https://sipoly.si/claim/ml-survey-2016-bestseller-2049/ ### AI will outperform humans at working in retail by 2031 (median) - Claimant: ML researchers (2016 expert survey) - Said: May 24, 2017 - Quote: "Researchers predict AI will outperform humans in many activities in the next ten years, such as translating languages (by 2024), writing high-school essays (by 2026), driving a truck (by 2027), working in retail (by 2031), writing a bestselling book (by 2049), and working as a surgeon (by 2053)." - Source: https://arxiv.org/abs/1705.08807 - Resolution criteria (our interpretation): Our interpretation: by end-2031, AI/robotic systems can do the full job of a retail salesperson at least as well as humans (deployed beyond pilots). - Resolve by: 2031-12-31 - Verdict: Open - URL: https://sipoly.si/claim/ml-survey-2016-retail-2031/ ### AI will outperform humans at driving a truck by 2027 (median) - Claimant: ML researchers (2016 expert survey) - Said: May 24, 2017 - Quote: "Researchers predict AI will outperform humans in many activities in the next ten years, such as translating languages (by 2024), writing high-school essays (by 2026), driving a truck (by 2027), working in retail (by 2031), writing a bestselling book (by 2049), and working as a surgeon (by 2053)." - Source: https://arxiv.org/abs/1705.08807 - Resolution criteria (our interpretation): Our interpretation: by end-2027, driverless trucks operate commercially at scale on public roads with safety at or better than human drivers. - Resolve by: 2027-12-31 - Verdict: Open - URL: https://sipoly.si/claim/ml-survey-2016-truck-2027/ ### AI will outperform humans at writing high-school essays by 2026 (median) - Claimant: ML researchers (2016 expert survey) - Said: May 24, 2017 - Quote: "Researchers predict AI will outperform humans in many activities in the next ten years, such as translating languages (by 2024), writing high-school essays (by 2026), driving a truck (by 2027), working in retail (by 2031), writing a bestselling book (by 2049), and working as a surgeon (by 2053)." - Source: https://arxiv.org/abs/1705.08807 - Resolution criteria (our interpretation): Our interpretation: by end-2026, AI-written high-school essays reliably score at or above typical human students in blind grading. - Resolve by: 2026-12-31 - Verdict: Open - URL: https://sipoly.si/claim/ml-survey-2016-hs-essays-2026/ ### Fully autonomous LA-to-New York demo drive "by the end of next year" - Claimant: Elon Musk - Said: Oct 19, 2016 - Quote: "we’ll be able to do a demonstration drive of full autonomy all the way from LA to New York, from home in LA to let’s say dropping you off in Time Square in New York, and then having the car go park itself, by the end of next year" - Source: https://techcrunch.com/2016/10/19/musk-targeting-coast-to-coast-test-drive-of-fully-self-driving-tesla-by-late-2017/ - Resolution criteria (our interpretation): A no-intervention coast-to-coast Tesla demonstration drive completed by 31 Dec 2017. - Resolve by: 2017-12-31 - Verdict: Wrong - URL: https://sipoly.si/claim/musk-coast-to-coast-2017/ ### Autonomous vehicles to provide the majority of Lyft rides within five years - Claimant: John Zimmer - Said: Sep 18, 2016 - Quote: "Autonomous vehicle fleets will quickly become widespread and will account for the majority of Lyft rides within 5 years." - Source: https://medium.com/@johnzimmer/the-third-transportation-revolution-27860f05fa91 - Resolution criteria (our interpretation): Majority of Lyft rides in autonomous vehicles by Sep 2021. - Resolve by: 2021-09-18 - Verdict: Wrong - URL: https://sipoly.si/claim/zimmer-lyft-majority-autonomous/ ### High-volume, fully autonomous (no steering wheel) ride-hailing vehicle in commercial operation in 2021 - Claimant: Ford Motor Company - Said: Aug 16, 2016 - Quote: "Ford today announces its intent to have a high-volume, fully autonomous SAE level 4-capable vehicle in commercial operation in 2021 in a ride-hailing or ride-sharing service." - Source: https://electrek.co/2016/08/16/ford-fully-autonomous-cars-high-volume-available-2021/ - Resolution criteria (our interpretation): Our interpretation: by 31 Dec 2021, Ford has a high-volume level-4 vehicle in commercial ride-hailing or ride-sharing service. - Resolve by: 2021-12-31 - Verdict: Broken - URL: https://sipoly.si/claim/ford-l4-ridehail-2021/ ### Deep learning to do "a lot better than radiologists" within five years - Claimant: Geoffrey Hinton - Said: 2016 - Quote: "People should stop training radiologists now. It’s just completely obvious that within five years, deep learning is going to do a lot better than radiologists" - Source: https://www.dotmed.com/news/story/39033 - Resolution criteria (our interpretation): Our interpretation: by end of 2021, deep learning systems outperform radiologists broadly enough that training new radiologists is no longer warranted. - Resolve by: 2021-12-31 - Verdict: Wrong - URL: https://sipoly.si/claim/hinton-radiologists-five-years/ ### Long Bet: no computer will pass the Turing Test by 2029 - Claimant: Mitch Kapor - Said: 2002 - Quote: "By 2029 no computer - or "machine intelligence" - will have passed the Turing Test." - Source: https://longbets.org/1/ - Resolution criteria (our interpretation): Our interpretation: as defined in the Long Bets terms: correct if no machine passes the Kapor-Kurzweil Turing Test protocol by the end of 2029. - Resolve by: 2029-12-31 - Verdict: Open - URL: https://sipoly.si/claim/kapor-no-turing-test-2029/ ### "Within thirty years, we will have the technological means to create superhuman intelligence" - Claimant: Vernor Vinge - Said: Mar 30, 1993 - Quote: "Within thirty years, we will have the technological means to create superhuman intelligence. Shortly after, the human era will be ended." - Source: https://edoras.sdsu.edu/~vinge/misc/singularity.html - Resolution criteria (our interpretation): Our interpretation: correct if by Mar 2023 humanity had the technological means to create AI (or augmented intelligence) that broadly exceeds human intelligence. - Resolve by: 2023-03-31 - Verdict: Wrong - URL: https://sipoly.si/claim/vinge-superhuman-intelligence-by-2023/