← Live scorecard
Explainer

How will we know when AGI has arrived?

Last updated: July 12, 2026 · Updated as verdicts change
By the AGI Scorecard team · methodology & independence
The test isn’t a benchmark score — it’s reliable autonomy. We’ll know AGI has arrived when a system can do the work of an AI researcher/engineer end-to-end, unsupervised — not when it clears one more eval. That single bar is why "AGI by 2027" is graded Open despite ~83% GDPval and ~80% SWE-Bench Pro: the capability is here, the autonomy isn’t.
62.5 Watch the bar move: AGI-2027 Thesis TrackerOne auditable score that changes the day the autonomy bar is cleared →

Three tests that get proposed — and which one counts

Proposed testProblem / status
Beat a benchmark (GDPval, SWE-Bench…)Largely passed — but scores ≠ autonomy
Pass as a "drop-in remote worker"The real bar — unmet: reliability lags
Autonomously conduct AI researchThe strictest bar — undemonstrated

The first is a moving target that keeps getting cleared without the world changing much. The serious tests are the last two: can the system reliably own a whole job end-to-end, and ultimately improve AI itself? Those are observable, falsifiable events — not vibes.

The concrete signals to watch

Why we track it as one number

Because "has AGI arrived?" fragments into eight specific, dated claims, the honest way to watch it is a scorecard, distilled into a single auditable figure. The AGI-2027 Thesis Tracker moves only when one of those verdicts changes — so the day the autonomy bar is actually cleared, the number moves, and you don’t have to parse the hype to notice. The headline claim resolves by January 1, 2028.

Frequently asked questions

How will we know when AGI has arrived?

When a system can reliably do the work of an AI researcher or engineer end-to-end and unsupervised — not when it clears another benchmark. As of mid-2026 that autonomy bar is unmet, which is why the AGI-by-2027 verdict is Open.

Is there a test for AGI?

The serious test isn't a single benchmark — it's reliable, accountable, long-horizon autonomy: owning a whole job end-to-end without human cleanup, and ultimately conducting AI research. Benchmark scores (GDPval ~83%, SWE-Bench ~80%) are necessary but not sufficient.

Why don't high benchmark scores mean AGI is here?

Because scoring well on scoped tasks isn't the same as reliably owning an entire job unsupervised. The gap between benchmark capability and accountable autonomy is exactly what still separates today's models from AGI.

When will we know if AGI by 2027 was right?

By January 1, 2028. The prediction is fulfilled if autonomous AI research is demonstrated by then, and Wrong if the deadline passes without it. The Thesis Tracker moves the moment that verdict changes.

The live scorecard updates as models ship and verdicts change.

View the live scorecard →

Get the weekly AGI progress briefing

We watch the autonomy bar so you don't have to. Subscribe to hear the moment it's cleared. Free, no hype.

Subscribe free →