Skip to content
AI IntelligenceAug 23, 2026AI Intelligence
Article

Prime Intellect published a benchmark evaluating 18 frontier models across 153 autonomous runs using the nanoGPT optimizer framework.

The study releases open traces of tool calls and reasoning scratchpads, providing developers with granular data on current model capabilities in self-directed tasks.

Data Cube AI EditorialSource: Prime Intellect
01

Source Brief

Prime Intellect published a benchmark evaluating 18 frontier models across 153 autonomous runs using the nanoGPT optimizer framework. The study releases open traces of tool calls and reasoning scratchpads, providing developers with granular data on current model capabilities in self-directed tasks.