← Live scorecard
Explainer

What is SWE-Bench? The agentic coding benchmark, explained

Last updated: July 10, 2026 · Updated as verdicts change
By the AGI Scorecard team · methodology & independence
The benchmark for real software-engineering work. SWE-Bench evaluates whether an AI agent can resolve real GitHub issues in real codebases — find the bug, write the patch, pass the tests. On this scorecard, top models reach roughly 80% on SWE-Bench Pro, the strongest single piece of evidence that agentic coding has crossed from demo to production capability.

What SWE-Bench measures

Unlike puzzle-style coding tests, SWE-Bench starts from actual issues filed against real open-source repositories. The agent gets the codebase and the issue text, and must produce a patch that resolves the issue and passes the project's own tests. That makes it an end-to-end test of the skills a working software engineer uses: navigating unfamiliar code, localizing a fault, editing multiple files, and verifying the fix. The harder Pro variant tightens the setup against contamination and inflated scores.

Why it matters for the AGI question

Situational Awareness treats coding as the leading indicator: AI that can do real software engineering is the on-ramp to AI that can automate AI research itself — the essay's definition of AGI. That is why this scorecard leans on SWE-Bench when grading the AGI-by-2027 prediction: at ~80% on SWE-Bench Pro, agentic coding is strong, and it is the main reason that verdict reads "trend intact" rather than "falling behind."

What ~80% does and doesn't mean

SignalReading
Resolving real, scoped issuesStrong (~80%)
Leading indicator for the 2027 thesisTrend intact
Unsupervised end-to-end engineeringNot demonstrated

A scoped GitHub issue is a well-posed task: the problem is described, the repo is given, the tests exist. Real engineering also involves deciding what to build, handling ambiguity, and owning a system over months. That gap is why the replace-programmers verdict reads "partly and unevenly" — and why autonomous AI research, the actual AGI bar, remains undemonstrated.

SWE-Bench and GDPval: two halves of one picture

SWE-Bench covers engineering; GDPval (~83%) covers broader knowledge work. Together they support the scorecard's capability verdict: models perform near the top of the skilled-human range on well-scoped professional tasks, while reliability without supervision — the thing that would flip the bigger verdicts — still lags.

Frequently asked questions

What is SWE-Bench?

SWE-Bench is a benchmark that tests whether AI agents can resolve real GitHub issues in real codebases — producing patches that pass the project's own tests. SWE-Bench Pro is the harder, contamination-resistant variant.

What do AI models score on SWE-Bench?

On this scorecard, top models reach roughly 80% on SWE-Bench Pro as of mid-2026 — strong enough that agentic coding is treated as production capability rather than a demo.

Does 80% on SWE-Bench mean AI can replace programmers?

Not by itself. Resolving scoped issues is a large part of engineering, but deciding what to build, handling ambiguity, and owning systems over time remain human work — which is why the scorecard grades the programmer question 'partly and unevenly.'

How does SWE-Bench relate to AGI?

Situational Awareness treats software engineering as the on-ramp to automating AI research — its definition of AGI. Strong SWE-Bench results keep the AGI-by-2027 trend intact, but the autonomous-research bar itself is still undemonstrated.

The live scorecard updates as models ship and verdicts change.

View the live scorecard →

Get the weekly AGI progress briefing

Verdict changes, lab milestones, and what they mean for the 2027 clock. Free — no hype, just signal.

Subscribe free →