OpenAI reports that adjusting two API settings substantially improved GPT-5.6’s performance on the ARC-AGI-3 benchmark. According to the description, retaining reasoning and enabling compaction tripled the model’s scores while also improving efficiency.

Why it matters

The account suggests that configuration choices at the API level can meaningfully affect benchmark outcomes, indicating that how a model is used may influence measured performance in addition to the model itself.

Who should care

Developers and teams working with GPT-5.6 or evaluating models against benchmarks like ARC-AGI-3 may find the described settings relevant when configuring their own workflows.