AI IntelligenceJul 8, 2026AI Intelligence
Article
OpenAI finds that about 30% of tasks in the SWE-Bench Pro benchmark are broken.
The company retracts its earlier recommendation to adopt the benchmark and urges caution when evaluating AI coding capabilities.
AI-generated: summaries written by AI from the linked sources. How we use AI
AI-generatedSource: Techmeme
01
Source Brief
OpenAI finds that about 30% of tasks in the SWE-Bench Pro benchmark are broken. The company retracts its earlier recommendation to adopt the benchmark and urges caution when evaluating AI coding capabilities.
02