Sign InOpen Brain
Back
OpenAIOfficial ReleaseOfficial SourcebenchmarkNew

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

Record reasoning retention and compaction with the model name because runtime settings can materially alter agent eval results.

OpenAIJul 29, 20262 min
Open SourceOpen MarkdownOpen JSON
Source Summary

Two API settings—reasoning retention and compaction—reportedly tripled GPT-5.6’s ARC-AGI-3 score. Agent evals should treat runtime configuration as part of the tested system.

Practical Implication

Record these settings alongside the model name in agent evaluations. Configuration can materially affect results, so defaults and explicit settings should not be compared as equivalent systems.

Agent-Ready Context
OpenAI says enabling **reasoning retention** and **compaction** produced **3× ARC-AGI-3 scores** for GPT-5.6 while also improving efficiency.

Record these settings alongside the model name in agent evaluations. Configuration can materially affect results, so defaults and explicit settings should not be compared as equivalent systems.

The supplied material provides no absolute scores, token usage, latency, or experimental detail, leaving the size and generality of the efficiency gain unclear.
Context Map
benchmark#agent-evals#benchmark-integrity#context-cachingGeneric AgentPrepare Coding Session
Rate This Item
Personal Note

No note yet. Notes are included in exported bundles.

Related — Every Edge Explained

No approved edges yet for this post.