How fast is AI improving?
Two speeds: inputs vs. the finish line
| What's improving | Rate (mid-2026) |
|---|---|
| Effective compute | ~0.5 OOM/yr — on trend |
| Benchmark scores | Fast — into the 80s% |
| Agentic / tool use | Fast — real production use |
| Reliable unsupervised autonomy | Slow — the bottleneck |
Measured by inputs and benchmarks, AI is improving very fast and roughly on the trend the optimists described. Measured by the one output that defines AGI — a system reliably owning a job end-to-end without supervision — progress is slower and harder to see, because benchmarks stopped being the binding constraint.
Why "is it slowing down?" gets the wrong answer
Skeptics point at plateauing benchmark headlines and say progress is stalling; boosters point at agentic demos and say it's accelerating. Both are looking at the fast axis. The honest read is that the inputs keep compounding (see are scaling laws dead?) while the decisive output — autonomy — is the slow, rate-limiting step. That split is exactly why the headline AGI-2027 verdict is Open, not On track.
One number for the rate that matters
Instead of arguing about vibes, watch the rate on the axis that changes the answer. The AGI-2027 Thesis Tracker distills whether the decisive capability is actually moving into one auditable score — it climbs only when a verdict changes, so a flat score is itself the signal that the slow axis hasn't broken yet. Resolves by January 1, 2028.
Frequently asked questions
Fast on inputs, uneven on the output that matters. Effective compute is scaling ~0.5 orders of magnitude per year and benchmark scores are into the 80s% (GDPval ~83%, SWE-Bench Pro ~80%), but reliable autonomous work — the AGI-defining capability — is improving much more slowly.
Not on the inputs — compute and benchmarks keep compounding roughly on trend. But the decisive output, reliable unsupervised autonomy, is the slow rate-limiting step. So 'slowing down' is true for the finish line and false for the engines driving toward it.
On agentic capability and benchmarks, yes. On the autonomy that defines AGI, no — that's the bottleneck. The two answers come from measuring different axes, which is why headlines disagree.
Watch the axis that changes the answer — reliable autonomous capability — rather than benchmark headlines. The AGI Scorecard's Thesis Tracker distills that into one auditable score that moves only when a verdict changes.
The live scorecard updates as models ship and verdicts change.
View the live scorecard →Get the weekly AGI progress briefing
We track the rate on the axis that actually moves the answer. Subscribe for the weekly read. Free, no hype.
Subscribe free →