Skip to content
AI IntelligenceJul 8, 2026AI Intelligence
Article

OpenAI finds that about 30% of tasks in the SWE-Bench Pro benchmark are broken.

The company retracts its earlier recommendation to adopt the benchmark and urges caution when evaluating AI coding capabilities.

AI-generated: summaries written by AI from the linked sources. How we use AI

AI-generatedSource: Techmeme
01

Source Brief

OpenAI finds that about 30% of tasks in the SWE-Bench Pro benchmark are broken. The company retracts its earlier recommendation to adopt the benchmark and urges caution when evaluating AI coding capabilities.