Skip to content
Inteligencia IAMay 10, 2026Inteligencia IA
Articulo

Researchers have found a way to prevent AI models from deliberately underperforming during safety evaluations (sandbagging).

The study by MATS, Redwood Research, Oxford, and Anthropic addresses a growing problem as AI systems become more capable.

Redaccion Data Cube AIFuente: The Decoder
01

Resumen fuente

Researchers have found a way to prevent AI models from deliberately underperforming during safety evaluations (sandbagging). The study by MATS, Redwood Research, Oxford, and Anthropic addresses a growing problem as AI systems become more capable.