Can AI replace knowledge workers? (the 2025/26 call, graded)
The prediction
One of the essay's most concrete near-term claims: by 2025/26, frontier models would reach or exceed the level of a smart college graduate across knowledge work — the precursor step to a "drop-in remote worker."
The evidence
The benchmark picture supports it. Frontier models score around 83% on knowledge-work benchmarks like GDPval, and agentic coding sits near 80% on SWE-Bench Pro, with agents deployed in real production workflows. On the benchmark bar, this is comfortably On track.
The caveat — and the flip condition
Benchmarks aren't the same as reliable autonomy. The pre-registered condition: this moves to Wrong if, by end-2026, frontier models still can't complete a majority of representative entry-level knowledge-work tasks end-to-end without human cleanup — the "drop-in coworker" bar, not the benchmark bar. The gap between scoring 80% on a test and reliably doing a junior employee's job unsupervised is exactly where the verdict could still turn — and with the flip condition dated end-2026, this is the scorecard's next verdict to resolve.
So can AI replace knowledge workers?
Partially, and unevenly. It clears the capability bar Aschenbrenner named and is automating slices of knowledge work; it has not become a reliable, unsupervised replacement for a knowledge worker. On track — with the reliability gap as the thing to watch.
Frequently asked questions
Partially. Frontier models clear the capability bar Aschenbrenner predicted for 2025/26 (~83% GDPval, ~80% SWE-Bench Pro) and automate slices of knowledge work, but aren't yet reliable unsupervised replacements. The prediction is graded On track.
On the benchmark bar, yes — graded On track. Frontier models reach ~83% on knowledge-work benchmarks and ~80% on agentic coding. The open question is 'drop-in coworker' reliability.
It flips to Wrong if, by end-2026, frontier models still can't complete a majority of representative entry-level knowledge-work tasks end-to-end without human cleanup.
The live scorecard updates as models ship and verdicts change.
View the live scorecard →Get the weekly AGI progress briefing
Verdict changes, lab milestones, and what they mean for the 2027 clock. Free — no hype, just signal.
Subscribe free →