What is SWE-Bench? The agentic coding benchmark, explained
What SWE-Bench measures
Unlike puzzle-style coding tests, SWE-Bench starts from actual issues filed against real open-source repositories. The agent gets the codebase and the issue text, and must produce a patch that resolves the issue and passes the project's own tests. That makes it an end-to-end test of the skills a working software engineer uses: navigating unfamiliar code, localizing a fault, editing multiple files, and verifying the fix. The harder Pro variant tightens the setup against contamination and inflated scores.
Why it matters for the AGI question
Situational Awareness treats coding as the leading indicator: AI that can do real software engineering is the on-ramp to AI that can automate AI research itself — the essay's definition of AGI. That is why this scorecard leans on SWE-Bench when grading the AGI-by-2027 prediction: at ~80% on SWE-Bench Pro, agentic coding is strong, and it is the main reason that verdict reads "trend intact" rather than "falling behind."
What ~80% does and doesn't mean
| Signal | Reading |
|---|---|
| Resolving real, scoped issues | Strong (~80%) |
| Leading indicator for the 2027 thesis | Trend intact |
| Unsupervised end-to-end engineering | Not demonstrated |
A scoped GitHub issue is a well-posed task: the problem is described, the repo is given, the tests exist. Real engineering also involves deciding what to build, handling ambiguity, and owning a system over months. That gap is why the replace-programmers verdict reads "partly and unevenly" — and why autonomous AI research, the actual AGI bar, remains undemonstrated.
SWE-Bench and GDPval: two halves of one picture
SWE-Bench covers engineering; GDPval (~83%) covers broader knowledge work. Together they support the scorecard's capability verdict: models perform near the top of the skilled-human range on well-scoped professional tasks, while reliability without supervision — the thing that would flip the bigger verdicts — still lags.
Frequently asked questions
SWE-Bench is a benchmark that tests whether AI agents can resolve real GitHub issues in real codebases — producing patches that pass the project's own tests. SWE-Bench Pro is the harder, contamination-resistant variant.
On this scorecard, top models reach roughly 80% on SWE-Bench Pro as of mid-2026 — strong enough that agentic coding is treated as production capability rather than a demo.
Not by itself. Resolving scoped issues is a large part of engineering, but deciding what to build, handling ambiguity, and owning systems over time remain human work — which is why the scorecard grades the programmer question 'partly and unevenly.'
Situational Awareness treats software engineering as the on-ramp to automating AI research — its definition of AGI. Strong SWE-Bench results keep the AGI-by-2027 trend intact, but the autonomous-research bar itself is still undemonstrated.
The live scorecard updates as models ship and verdicts change.
View the live scorecard →Get the weekly AGI progress briefing
Verdict changes, lab milestones, and what they mean for the 2027 clock. Free — no hype, just signal.
Subscribe free →