OpenAI Flags Major Flaws in SWE-Bench Pro

OpenAI's audit reveals approximately 30% of SWE-Bench Pro's coding tasks are flawed, prompting the company to retract its recommendation for the benchmark.

A stylized depiction of AI code review, with glowing lines of code being analyzed by an abstract AI entity.
OpenAI's audit reveals significant flaws in the SWE-Bench Pro coding benchmark, impacting AI model evaluations.· OpenAI News
Visual TL;DR
Previous Benchmark IssuesDriver
SWE-bench Verified had fundamental design and contamination problems identified earlier
From the article 2 mentionsOpenAI has identified significant issues within SWE-Bench Pro, a prominent benchmark for evaluating AI coding agents, estimating that roughly 30% of its tasks are broken.
OpenAI Audit MethodologyCore
comprehensive datapoint analysis pipeline reviewed model attempts and failure traces
From the article 2 mentionsOpenAI's methodology for this SWE-Bench Pro audit involved a comprehensive datapoint analysis pipeline.
SWE-Bench Pro FlawsDriver
From the article 5 mentionsOpenAI has identified significant issues within SWE-Bench Pro, a prominent benchmark for evaluating AI coding agents, estimating that roughly 30% of its tasks are broken.
Human Engineer ReviewCore
From the articleEach flagged task then underwent scrutiny through multiple investigator-agent passes and independent review by five experienced software engineers.
200 Tasks BrokenOutcome
From the article 2 mentionsOf the 731 public split tasks, the analysis pipeline flagged 200 (27.4%) as broken, while human annotation identified 249 (34.1%).
OpenAI Retracts RecommendationOutcome
company previously encouraged adoption, now withdraws its support for the benchmark
From the articleGiven the widespread SWE-Bench Pro task issues, OpenAI has retracted its earlier recommendation for the benchmark.
Challenges in AI EvalContext
From the articleThis detailed audit, published by OpenAI News, underscores the challenges in creating reliable coding evaluations for advanced AI models.

OpenAI has identified significant issues within SWE-Bench Pro, a prominent benchmark for evaluating AI coding agents, estimating that roughly 30% of its tasks are broken. This detailed audit, published by OpenAI News, underscores the challenges in creating reliable coding evaluations for advanced AI models.

StartupHub data

Companies working on this

Profiles of the companies named in this story, with funding and a one-liner from our database.

OpenAI
Private / $100B+ est
OpenAI is an AI research and deployment company dedicated to ensuring that artificial general intelligence benefits all of humanity.

The company previously encouraged the community to adopt SWE-Bench Pro after finding fundamental design and contamination problems in its predecessor, SWE-bench Verified. However, their latest investigation reveals similar concerns.

OpenAI's methodology for this SWE-Bench Pro audit involved a comprehensive datapoint analysis pipeline. This system reviewed model attempts, task metadata, and failure traces to flag potential evaluation flaws. Each flagged task then underwent scrutiny through multiple investigator-agent passes and independent review by five experienced software engineers.

Of the 731 public split tasks, the analysis pipeline flagged 200 (27.4%) as broken, while human annotation identified 249 (34.1%). These coding benchmark evaluation flaws fall into four primary categories:

  • Overly strict tests: These enforce specific implementation details not explicitly stated in the prompt, invalidating functionally correct submissions.
  • Underspecified prompts: Key requirements are omitted, enforced only by hidden tests and not reasonably inferable.
  • Low-coverage tests: These inadequately check the requested feature, allowing incomplete fixes to pass.
  • Misleading prompts: Models are directed toward incorrect behavior or given instructions that contradict test requirements.

These findings highlight the ongoing difficulty in curating fair yet challenging benchmarks, especially as AI model capabilities improve. The audit suggests a growing utility for agents in scalable data quality checks, helping to surface issues that were once impractical to find at scale.

Given the widespread SWE-Bench Pro task issues, OpenAI has retracted its earlier recommendation for the benchmark. The company advises developers to carefully examine results derived from SWE-Bench Pro, emphasizing that valid and informative evaluations are crucial for sound deployment and safety decisions under OpenAI’s Preparedness Framework.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.

More from Daniel Singer