Sign InOpen Brain
OpenAIOfficial ReleaseOfficial Source

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

Two API settings—reasoning retention and compaction—reportedly tripled GPT-5.6’s ARC-AGI-3 score. Agent evals should treat runtime configuration as part of the tested system.

OpenAI · Jul 29, 2026
Open Source Open MarkdownOpen JSON
Source Summary

OpenAI says enabling **reasoning retention** and **compaction** produced **3× ARC-AGI-3 scores** for GPT-5.6 while also improving efficiency.

Practical Implication

Record these settings alongside the model name in agent evaluations. Configuration can materially affect results, so defaults and explicit settings should not be compared as equivalent systems.

Agent-Ready Context
OpenAI says enabling **reasoning retention** and **compaction** produced **3× ARC-AGI-3 scores** for GPT-5.6 while also improving efficiency.

Record these settings alongside the model name in agent evaluations. Configuration can materially affect results, so defaults and explicit settings should not be compared as equivalent systems.

The supplied material provides no absolute scores, token usage, latency, or experimental detail, leaving the size and generality of the efficiency gain unclear.
Connected Context · Feed7 Judgment

This provides direct evidence that agent configuration can dominate a benchmark result: the same named model scored three times higher with reasoning retention and compaction enabled. It strengthens the case for treating model, harness, state-management settings, and resource conditions as one evaluated system, while absent absolute scores and experimental details limit generalization beyond ARC-AGI-3.

Context Map
benchmark#agent-evals#benchmark-integrity#context-caching
Uncertainty
The supplied material provides no absolute scores, token usage, latency, or experimental detail, leaving the size and generality of the efficiency gain unclear.