How will we know when AGI has arrived?
Three tests that get proposed — and which one counts
| Proposed test | Problem / status |
|---|---|
| Beat a benchmark (GDPval, SWE-Bench…) | Largely passed — but scores ≠ autonomy |
| Pass as a "drop-in remote worker" | The real bar — unmet: reliability lags |
| Autonomously conduct AI research | The strictest bar — undemonstrated |
The first is a moving target that keeps getting cleared without the world changing much. The serious tests are the last two: can the system reliably own a whole job end-to-end, and ultimately improve AI itself? Those are observable, falsifiable events — not vibes.
The concrete signals to watch
- A model runs a real, multi-week project unsupervised and is accountable for the outcome — not just assisting a human who cleans up.
- Frontier labs report AI meaningfully automating their own research engineering (the intelligence-explosion trigger).
- Adoption shifts from "assist" to "replace" for whole roles, not tasks — the signal behind the jobs question.
Why we track it as one number
Because "has AGI arrived?" fragments into eight specific, dated claims, the honest way to watch it is a scorecard, distilled into a single auditable figure. The AGI-2027 Thesis Tracker moves only when one of those verdicts changes — so the day the autonomy bar is actually cleared, the number moves, and you don’t have to parse the hype to notice. The headline claim resolves by January 1, 2028.
Frequently asked questions
When a system can reliably do the work of an AI researcher or engineer end-to-end and unsupervised — not when it clears another benchmark. As of mid-2026 that autonomy bar is unmet, which is why the AGI-by-2027 verdict is Open.
The serious test isn't a single benchmark — it's reliable, accountable, long-horizon autonomy: owning a whole job end-to-end without human cleanup, and ultimately conducting AI research. Benchmark scores (GDPval ~83%, SWE-Bench ~80%) are necessary but not sufficient.
Because scoring well on scoped tasks isn't the same as reliably owning an entire job unsupervised. The gap between benchmark capability and accountable autonomy is exactly what still separates today's models from AGI.
By January 1, 2028. The prediction is fulfilled if autonomous AI research is demonstrated by then, and Wrong if the deadline passes without it. The Thesis Tracker moves the moment that verdict changes.
The live scorecard updates as models ship and verdicts change.
View the live scorecard →Get the weekly AGI progress briefing
We watch the autonomy bar so you don't have to. Subscribe to hear the moment it's cleared. Free, no hype.
Subscribe free →