What AI coding benchmarks actually measure, and where they mislead

SWE-bench, SWE-bench Verified and LiveCodeBench are the scores labs quote to prove coding ability. They measure something real and narrow, and contamination, weak tests and saturation mean a high number says less about engineering reliability than it looks.

By Yash Malviya

Published

A person is typing code on a laptop, focusing on the screen with programming script
Photo: Lukas Blazek / Pexels

The claim and the catch Every few weeks a lab reports a new high on SWE-bench Verified, the most cited test of whether AI can "do software engineering". The numbers climb fast, and the headline that follows is usually that models are closing in on human engineers. Before accepting that, it is worth asking a plain question: what do these benchmarks actually measure, and what does a high score leave out?

The honest answer is that the main code benchmarks measure something real and narrow. They check whether a model can produce a patch that makes a specific set of unit tests pass, on a specific kind of task, drawn from a specific slice of open-source Python. That is useful. It is not the same as reliable engineering, and the gap between the two is where most of the confusion lives.

How the benchmarks are built SWE-bench, published by a Princeton-led team at ICLR 2024, collected 2,294 tasks by crawling merged pull requests and their linked issues across 12 popular Python repositories such as Django, scikit-learn and SymPy. Each task hands the model a bug report and the codebase at that point in time, then asks for a patch. The patch is judged by running the project's own tests: the ones that failed before the fix and should pass after it.

Because many of the original tasks turned out to be ambiguous or impossible, OpenAI worked with the SWE-bench authors to build SWE-bench Verified, a cleaned subset. According to Epoch AI, 93 professional developers screened 1,699 samples, with three annotators per sample, to produce 500 tasks judged clear and solvable. The creators estimated the original set carried a 5 to 10 percent error rate. Verified is now the default scoreboard, and that is the number labs quote.

LiveCodeBench takes a different route. Instead of repository bugs it pulls competitive-programming problems from LeetCode, AtCoder and Codeforces, 511 of them released between May 2023 and May 2024, and tags each with its publication date.

“A model that scores 70 percent on SWE-bench Verified has not fixed 70 percent of real bugs in any general sense.”

A digital tablet showing a web analytics dashboard with graphs and charts
AI coding benchmarks reduce messy engineering work to pass-or-fail unit tests on curated tasks. Photo: weCare Media / Pexels

The contamination problem That date tag is the whole point. A model trained on the public internet may have already seen a benchmark's problems and their solutions, so a high score can reflect memory rather than skill. LiveCodeBench's authors showed the pattern directly: one model's pass rate fell from roughly 60 to near zero on problems published after its release date, exactly where memorisation would stop helping. By scoring models only on problems released after their training cutoff, LiveCodeBench tries to measure reasoning rather than recall.

SWE-bench is more exposed to this. An independent study, SWE-Bench+ by Aleithan and colleagues in 2024, found that more than 94 percent of its issues predate the knowledge cutoffs of the models being tested. Worse, 32.67 percent of the patches that models got "right" had the solution sitting in plain view in the issue report or its comments. The model was not solving the bug so much as copying the answer. Another 31.08 percent passed only because the tests were too weak to catch a wrong fix. When the researchers filtered these cases out, one well-known agent's resolution rate collapsed from 12.47 percent to 3.97 percent.

What a high score does and does not say Put those findings together and a benchmark score starts to look like a compound of several things: genuine problem-solving, memorised answers, leaked hints, and the luck of a weak test suite. A model that scores 70 percent on SWE-bench Verified has not fixed 70 percent of real bugs in any general sense. It has produced patches that satisfy a particular test harness on a curated set of Python issues, some of which it may have seen before.

What the score does say is still worth something. It is a reproducible, test-grounded signal that tracks how useful a model is inside a coding agent, and it is far better than a demo chosen to impress. But it is narrow in ways that matter for anyone deciding whether to trust a model in production. The tasks are self-contained and short. They arrive with a known-good test that defines success. Real engineering rarely offers that. Requirements are vague, the right behaviour is contested, the blast radius of a change is unknown, and no oracle tells you when you are done. A benchmark that supplies a clear issue and a ready-made test removes exactly the parts of the job that make it hard.

Narrowness shows up in coverage too. Both SWE-bench and LiveCodeBench lean heavily on Python and on domains that are well represented online. They say little about large legacy codebases, concurrency bugs, performance regressions, security review, or the long multi-file changes that fill a real sprint.

Saturation and overfitting The newest problem is success itself. Frontier models now report scores above 70 percent on SWE-bench Verified, and several public leaderboards put the top of the table above 90 percent, leaving only a few points of headroom. When a benchmark is nearly solved it stops separating models, and the pressure to optimise directly for it grows. Harder successors such as SWE-bench Pro, where reported scores sit far lower, exist precisely because the Verified numbers have bunched near the ceiling.

Overfitting is the quieter risk. Once a benchmark is the metric everyone chases, teams tune agents, prompts and training data toward its particular shape. That can lift the score without lifting real-world reliability, especially given the leakage and weak-test issues already documented. Living benchmarks that keep adding fresh, post-cutoff tasks are one answer, because a target that moves is harder to overfit.

How to read the numbers None of this makes the benchmarks worthless. It means they should be read as what they are: a lower bound on a narrow, well-specified slice of coding, not a verdict on engineering judgement. A careful reader treats a high SWE-bench score as necessary but not sufficient, asks which version and which date window produced it, and weighs it against contamination-aware tests and the team's own private evaluations. The most reliable signal is still the oldest one. Watch how a model performs on problems it could not have seen, inside the messy work you actually do.

Frequently asked questions

What is the difference between SWE-bench and SWE-bench Verified?

SWE-bench is the original set of 2,294 GitHub issue tasks from 12 Python repositories. SWE-bench Verified is a 500-task subset that OpenAI and the authors cleaned with 93 developers, after estimating a 5 to 10 percent error rate in the original (Epoch AI, 2024).

Is SWE-bench contaminated?

It is exposed to contamination. The SWE-Bench+ study found more than 94 percent of its issues predate model knowledge cutoffs, 32.67 percent of 'solved' patches had the solution leaked in the issue text, and 31.08 percent passed on weak tests (Aleithan et al., 2024).

How does LiveCodeBench avoid contamination?

It tags each contest problem with a release date and scores models only on problems published after their training cutoff, so memorised answers cannot help. Authors showed one model's pass rate fell from about 60 to near zero on newer problems (LiveCodeBench, 2024).

Does a high SWE-bench score mean a model is a good engineer?

Not on its own. The score reflects narrow, test-graded Python tasks and is affected by leakage, weak tests and saturation. It is a useful lower bound, not a verdict on real-world engineering reliability.

Sources

  1. 1

    SWE-bench Verified, Epoch AI (13 August 2024)