Why we no longer evaluate SWE-bench Verified
OpenAI discontinues SWE-bench Verified evaluations due to contamination and inaccuracies, shifting focus to SWE-bench Pro.
OpenAI has officially announced the discontinuation of evaluations for SWE-bench Verified, a benchmark designed to assess coding progress in software engineering. This decision stems from significant concerns regarding the reliability of the tests, which have been found to contain flaws and inaccuracies that compromise their effectiveness. The organization cited issues such as training leakage and contaminated data as primary reasons for this shift, indicating that these factors have led to misleading results that do not accurately reflect coding capabilities or progress.
As a response to these challenges, OpenAI is transitioning to SWE-bench Pro, a more reliable alternative that promises to deliver a more accurate evaluation of coding skills. This new benchmark aims to address the shortcomings identified in SWE-bench Verified, providing a clearer and more trustworthy measure of software engineering proficiency. The move is part of OpenAI's broader commitment to ensuring that its evaluation tools maintain high standards of integrity and accuracy, which are crucial for developers and researchers relying on these assessments.
Key facts
| Field | Detail |
|---|---|
| Discontinued Benchmark | SWE-bench Verified |
| New Benchmark | SWE-bench Pro |
| Main Issues | Flawed tests, training leakage, inaccuracies |
| Reason for Discontinuation | Contamination and misleading results |
| Commitment | High standards of integrity and accuracy |
The decision to discontinue SWE-bench Verified reflects a growing awareness within the AI and software development communities about the importance of reliable evaluation metrics. In recent years, various benchmarks have come under scrutiny for their ability to accurately assess performance. For instance, similar concerns were raised about the GLUE benchmark in natural language processing, which led to the development of more robust alternatives. OpenAI's proactive approach in addressing these issues is indicative of a larger trend towards improving the quality of evaluation tools in the tech industry.
Looking ahead, the introduction of SWE-bench Pro is expected to set a new standard for coding assessments. As developers and researchers begin to adopt this new benchmark, it will be crucial to monitor its effectiveness in providing accurate evaluations. The transition not only aims to rectify past inaccuracies but also to establish a more reliable framework for measuring coding progress, which could influence how software engineering skills are assessed in the future. The success of SWE-bench Pro will likely depend on continuous feedback from the community and ongoing refinements to ensure it meets the evolving needs of the industry.
Source: OpenAI News · Read original →
Discussion
Comment here after signing in, or share the story to continue the conversation elsewhere.
Instagram & TikTok: copy the link and paste into a Story, Reel, or post caption.
Log in or create an account to comment — Google / GitHub / X when those providers are configured.
No comments yet — start the thread.

