What is GDPval? The knowledge-work benchmark, explained
What GDPval measures
GDPval is a benchmark built to test something narrower and more useful than trivia or puzzle-solving: can a model produce the deliverables of real, economically valuable knowledge work? Instead of abstract exam questions, it draws tasks from a spread of occupations — the reports, analyses, and documents that people are actually paid to create — and scores a model's output against expert-produced reference work. The name is a nod to GDP: the point is work that has economic value, not just benchmark points.
Why it matters for the AGI question
Aschenbrenner's Situational Awareness predicted that models would reach the level of a smart college graduate on knowledge work well before 2027. GDPval is one of the cleanest ways to test that claim, because it targets output that resembles a real job rather than a synthetic test. On this scorecard, top models reach roughly 83% on GDPval-style knowledge work — which is why the "AI outpaces knowledge workers" prediction grades On track.
What the ~83% does and doesn't tell you
| Signal | Reading |
|---|---|
| Raw capability on knowledge tasks | Strong (~83%) |
| Trend vs. the 2027 prediction | On track |
| Drop-in reliability (unsupervised) | Still lags |
A high GDPval score means a model can produce work comparable to a skilled human on a large fraction of tasks. It does not mean the model is a reliable, unsupervised drop-in worker — the gap between scoring well on a benchmark and being trusted to run a task end-to-end without review is exactly where the remaining AGI uncertainty lives. That distinction is why capability is graded On track while full autonomy is not.
How it fits the other benchmarks
GDPval is the knowledge-work counterpart to SWE-Bench Pro (~80%), which measures agentic software engineering. Together they paint a consistent picture: on task capability, models are already near the top of the human range in several white-collar domains; on autonomy and reliability, they still trail. That is the core tension the "can AI replace knowledge workers?" verdict tracks.
One benchmark is a data point. The Tracker is the whole picture.
GDPval feeds just one of eight verdicts. Our auditable AGI-2027 Thesis Tracker rolls them all into a single 0–100 score — currently 62.5/100. Subscribe to hear when it moves.
See the live Thesis Tracker →Frequently asked questions
GDPval is a benchmark that scores AI models on real, economically valuable knowledge-work tasks drawn from many occupations, comparing their output against expert-produced reference deliverables. It's designed to measure job-like work, not abstract test questions.
On this scorecard, top models reach roughly 83% on GDPval-style knowledge work — strong enough that the 'AI outpaces college-grad knowledge workers' prediction is graded On track.
Not on its own. A ~83% score shows strong task capability, but it doesn't prove reliable, unsupervised performance. Drop-in reliability still lags the benchmark, which is why full autonomy remains an open question.
GDPval measures general knowledge work; SWE-Bench Pro (~80%) measures agentic software engineering. Both show strong task capability but a remaining gap on unsupervised autonomy.
The live scorecard updates as models ship and verdicts change.
View the live scorecard →Get the weekly AGI progress briefing
Verdict changes, lab milestones, and what they mean for the 2027 clock. Free — no hype, just signal.
Subscribe free →