---
格式版本: 2
标题: "开启两项设置后，我们在 ARC-AGI-3 基准测试中的得分增至三倍 | OpenAI"
原文链接: "https://openai.com/zh-Hans-CN/index/how-two-settings-tripled-our-arc-agi-3-scores/"
发布日期: "2026-07-28"
发布时间校准状态: "provisional"
发布时间需复核: "是"
发布时间来源: "fallback:local:strict_original_body"
发布时间证据: "2026年7月28日"
发布时间校准原因: "LLM 仲裁服务异常，按现有强证据临时采用 2026-07-28，后续需复核"
发布时间校准置信度: "medium"
发布时间候选数量: 14
发布时间严格候选数量: 3
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-10T16:21:36+08:00"
发布时间仲裁状态: "service-error"
发布时间仲裁尝试次数: 2
发布时间仲裁耗时毫秒: 52850
发现时间: "2026-08-10T16:18:32+08:00"
入库时间: "2026-08-10T08:22:30.110Z"
来源平台: "固定入口"
搜索渠道: "fixed_url"
搜索词: "https://openai.com/news/"
匹配关键词:
  - "performance"
  - "计算"
  - "部署"
  - "性能"
相关厂家:
  - "OpenAI"
相关专家:
  []
内容类型: "网页"
抓取工具: "CDP Render"
清洗工具: "CDP Text + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 10
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "内容为GPT-5.6在ARC-AGI-3基准测试上的性能改进，讨论API设置，与超节点/AI Rack/机柜级AI基础设施完全无关。"
AI质检模型: "deepseek-v4-flash"
AI质检时间: "2026-08-10T16:22:53+08:00"
AI主题相关性: 2
AI来源权威性: 8
AI新颖性: 0
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 0
采集批次: "2026年8月10日15点37分56秒"
采集批次ID: "20260810-153756-703"
去重键: "https://openai.com/zh-Hans-CN/index/how-two-settings-tripled-our-arc-agi-3-scores"
---

2026年7月29日

[研究](https://openai.com/zh-Hans-CN/news/research/) [刊发](https://openai.com/zh-Hans-CN/research/index/publication/)

*这段加速视频展示了 GPT‑5.6 Sol 尝试解决 ARC-AGI-3 基准测试中的谜题：左侧使用官方执行框架，右侧使用我们保留推理并启用压缩的 Responses API 执行框架。在* [*这款游戏* ⁠](https://arcprize.org/tasks/cd82) *的排行榜上，没有任何前沿模型能通过第一关之后的关卡。使用我们的执行框架后，GPT‑5.6 Sol 通关了全部六关。*

最初看到 GPT‑5.6 Sol 在 [ARC-AGI-3⁠](https://arcprize.org/arc-agi/3) 基准测试中的低分时，我们十分困惑。

GPT‑5.6 Sol 已经解决了数学界长期悬而未决的难题，例如 [圈双覆盖猜想⁠](https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98d31/cdc_proof.pdf) ，也通关了《宝可梦 火红》等游戏。但在由二维益智游戏组成的 ARC-AGI-3 基准测试中，GPT‑5.6 Sol 仅得 7.8%，GPT‑5.5 更是几乎无法游玩，得分只有 0.4%。

二维益智游戏对我们的模型来说格外困难吗？还是另有原因？

基准测试很少只衡量 AI 模型本身。它们还会衡量一些不太显眼的选择，包括 API 设置、执行框架设计和提示方式。在 ARC-AGI-3 上，我们发现，开启 ChatGPT 和 Codex 中使用的两项 API 设置——保留推理和压缩——可使公开任务集的得分增至三倍，同时将输出词元降至六分之一。

GPT-5.6 Sol 在 ARC-AGI-3 公开集上的表现

<svg xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink" version="1.1" width="581" height="363" viewBox="0 0 581 363" style="background-color: transparent;"><g fill="none" stroke-miterlimit="10" transform="translate(76,30)"><g role="graphics-object" aria-roledescription="group mark container"><g transform="translate(0,0)"><g><g role="graphics-symbol" aria-roledescription="axis" aria-label="X-axis titled '每局输出词元数' for a linear scale with values from 0 to 3500000"><g transform="translate(0.5,268.5)"><g><g pointer-events="none"><line transform="translate(69,0)" x2="0" y2="5" stroke="currentColor" stroke-width="1" opacity="1"></line><line transform="translate(139,0)" x2="0" y2="5" stroke="currentColor" stroke-width="1" opacity="1"></line><line transform="translate(208,0)" x2="0" y2="5" stroke="currentColor" stroke-width="1" opacity="1"></line><line transform="translate(277,0)" x2="0" y2="5" stroke="currentColor" stroke-width="1" opacity="1"></line><line transform="translate(346,0)" x2="0" y2="5" stroke="currentColor" stroke-width="1" opacity="1"></line><line transform="translate(416,0)" x2="0" y2="5" stroke="currentColor" stroke-width="1" opacity="1"></line><line transform="translate(485,0)" x2="0" y2="5" stroke="currentColor" stroke-width="1" opacity="1"></line></g><g pointer-events="none"><text text-anchor="middle" transform="translate(69.28571428571428,24)" font-family="SF Mono, monospace" font-size="12px" font-weight="normal" fill="currentColor" opacity="1">0.5M</text> <text text-anchor="middle" transform="translate(138.57142857142856,24)" font-family="SF Mono, monospace" font-size="12px" font-weight="normal" fill="currentColor" opacity="1">1M</text> <text text-anchor="middle" transform="translate(207.85714285714283,24)" font-family="SF Mono, monospace" font-size="12px" font-weight="normal" fill="currentColor" opacity="1">1.5M</text> <text text-anchor="middle" transform="translate(277.1428571428571,24)" font-family="SF Mono, monospace" font-size="12px" font-weight="normal" fill="currentColor" opacity="1">2M</text> <text text-anchor="middle" transform="translate(346.42857142857144,24)" font-family="SF Mono, monospace" font-size="12px" font-weight="normal" fill="currentColor" opacity="1">2.5M</text> <text text-anchor="middle" transform="translate(415.71428571428567,24)" font-family="SF Mono, monospace" font-size="12px" font-weight="normal" fill="currentColor" opacity="1">3M</text> <text text-anchor="middle" transform="translate(485,24)" font-family="SF Mono, monospace" font-size="12px" font-weight="normal" fill="currentColor" opacity="1">3.5M</text></g> <g pointer-events="none"><line transform="translate(0,0)" x2="485" y2="0" stroke="currentColor" stroke-width="1" opacity="1"></line></g><g pointer-events="none"><text text-anchor="middle" transform="translate(242.5,57)" font-family="OpenAI Sans, OpenAI Sans Variable Scripts, sans-serif" font-size="14px" font-weight="normal" fill="currentColor" opacity="1">每局输出词元数</text></g></g></g></g> <g role="graphics-symbol" aria-roledescription="axis" aria-label="Y-axis titled '得分' for a linear scale with values from 0% to 40%"><g transform="translate(0.5,0.5)"><g><g pointer-events="none"><line transform="translate(0,268)" x2="-5" y2="0" stroke="currentColor" stroke-width="1" opacity="1"></line><line transform="translate(0,201)" x2="-5" y2="0" stroke="currentColor" stroke-width="1" opacity="1"></line><line transform="translate(0,134)" x2="-5" y2="0" stroke="currentColor" stroke-width="1" opacity="1"></line><line transform="translate(0,67)" x2="-5" y2="0" stroke="currentColor" stroke-width="1" opacity="1"></line><line transform="translate(0,0)" x2="-5" y2="0" stroke="currentColor" stroke-width="1" opacity="1"></line></g><g pointer-events="none"><text text-anchor="end" transform="translate(-15,272)" font-family="SF Mono, monospace" font-size="12px" font-weight="normal" fill="currentColor" opacity="1">0%</text> <text text-anchor="end" transform="translate(-15,205)" font-family="SF Mono, monospace" font-size="12px" font-weight="normal" fill="currentColor" opacity="1">10%</text> <text text-anchor="end" transform="translate(-15,138)" font-family="SF Mono, monospace" font-size="12px" font-weight="normal" fill="currentColor" opacity="1">20%</text> <text text-anchor="end" transform="translate(-15,71.00000000000003)" font-family="SF Mono, monospace" font-size="12px" font-weight="normal" fill="currentColor" opacity="1">30%</text> <text text-anchor="end" transform="translate(-15,4)" font-family="SF Mono, monospace" font-size="12px" font-weight="normal" fill="currentColor" opacity="1">40%</text></g> <g pointer-events="none"><line transform="translate(0,268)" x2="0" y2="-268" stroke="currentColor" stroke-width="1" opacity="1"></line></g><g pointer-events="none"><text text-anchor="middle" transform="translate(-56.673828125,134) rotate(-90) translate(0,-3)" font-family="OpenAI Sans, OpenAI Sans Variable Scripts, sans-serif" font-size="14px" font-weight="normal" fill="currentColor" opacity="1">得分</text></g></g></g></g> <g role="graphics-object" aria-roledescription="group mark container"><g transform="translate(0,0)"><g><g role="graphics-object" aria-roledescription="line mark container"><path aria-label="output_tokens_per_game: 84316; score: 4%; harness_label: Harness with retained reasoning and compaction; effort_order: 0; series: GPT-5.6 Sol · Harness with retained reasoning and compaction" role="graphics-symbol" aria-roledescription="line mark" d="M11.684,243.21L19.669,219.09L33.709,178.22L59.383,95.81L67.274,11.39" stroke="currentColor" stroke-width="1.5" stroke-dasharray="1,0"></path></g><g role="graphics-object" aria-roledescription="symbol mark container"><path aria-label="output_tokens_per_game: 84316; score: 4%; 模型: GPT-5.6 Sol; 执行框架: Harness with retained reasoning and compaction; 推理强度: Low; 得分: 3.7%; 每局输出词元数: 84,316; 输出词元总数: 2,107,901" role="graphics-symbol" aria-roledescription="point" transform="translate(11.683788571428572,243.20999999999998)" d="M4.183,0A4.183,4.183,0,1,1,-4.183,0A4.183,4.183,0,1,1,4.183,0" fill="currentColor" stroke-width="1.5" opacity="1"></path><path aria-label="output_tokens_per_game: 141939; score: 7%; 模型: GPT-5.6 Sol; 执行框架: Harness with retained reasoning and compaction; 推理强度: Medium; 得分: 7.3%; 每局输出词元数: 141,939; 输出词元总数: 3,548,473" role="graphics-symbol" aria-roledescription="point" transform="translate(19.66869,219.09)" d="M4.183,0A4.183,4.183,0,1,1,-4.183,0A4.183,4.183,0,1,1,4.183,0" fill="currentColor" stroke-width="1.5" opacity="1"></path><path aria-label="output_tokens_per_game: 243258; score: 13%; 模型: GPT-5.6 Sol; 执行框架: Harness with retained reasoning and compaction; 推理强度: High; 得分: 13.4%; 每局输出词元数: 243,258; 输出词元总数: 6,081,458" role="graphics-symbol" aria-roledescription="point" transform="translate(33.70860857142857,178.22)" d="M4.183,0A4.183,4.183,0,1,1,-4.183,0A4.183,4.183,0,1,1,4.183,0" fill="currentColor" stroke-width="1.5" opacity="1"></path><path aria-label="output_tokens_per_game: 428540; score: 26%; 模型: GPT-5.6 Sol; 执行框架: Harness with retained reasoning and compaction; 推理强度: Xhigh; 得分: 25.7%; 每局输出词元数: 428,540; 输出词元总数: 10,713,499" role="graphics-symbol" aria-roledescription="point" transform="translate(59.383399999999995,95.81000000000002)" d="M4.183,0A4.183,4.183,0,1,1,-4.183,0A4.183,4.183,0,1,1,4.183,0" fill="currentColor" stroke-width="1.5" opacity="1"></path><path aria-label="output_tokens_per_game: 485485; score: 38%; 模型: GPT-5.6 Sol; 执行框架: Harness with retained reasoning and compaction; 推理强度: Max; 得分: 38.3%; 每局输出词元数: 485,485; 输出词元总数: 12,137,135" role="graphics-symbol" aria-roledescription="point" transform="translate(67.27435,11.389999999999995)" d="M4.183,0A4.183,4.183,0,1,1,-4.183,0A4.183,4.183,0,1,1,4.183,0" fill="currentColor" stroke-width="1.5" opacity="1"></path></g></g></g><g transform="translate(0,0)"><g><g role="graphics-object" aria-roledescription="line mark container"><path aria-label="output_tokens_per_game: 58809; score: 1%; harness_label: Official harness; effort_order: 0; series: GPT-5.6 Sol · Official harness" role="graphics-symbol" aria-roledescription="line mark" d="M8.149,261.97L28.71,257.95L100.906,233.16L178.119,219.76L401.995,178.89" stroke="currentColor" stroke-width="1.5" stroke-dasharray="8,5"></path></g><g role="graphics-object" aria-roledescription="symbol mark container"><path aria-label="output_tokens_per_game: 58809; score: 1%; 模型: GPT-5.6 Sol; 执行框架: Official harness; 推理强度: Low; 得分: 0.9%; 每局输出词元数: 58,809; 输出词元总数: 1,470,234" role="graphics-symbol" aria-roledescription="point" transform="translate(8.149247142857142,261.97)" d="M4.183,0A4.183,4.183,0,1,1,-4.183,0A4.183,4.183,0,1,1,4.183,0" fill="currentColor" stroke-width="1.5" opacity="1"></path><path aria-label="output_tokens_per_game: 207188; score: 2%; 模型: GPT-5.6 Sol; 执行框架: Official harness; 推理强度: Medium; 得分: 1.5%; 每局输出词元数: 207,188; 输出词元总数: 5,179,705" role="graphics-symbol" aria-roledescription="point" transform="translate(28.710337142857146,257.95)" d="M4.183,0A4.183,4.183,0,1,1,-4.183,0A4.183,4.183,0,1,1,4.183,0" fill="currentColor" stroke-width="1.5" opacity="1"></path><path aria-label="output_tokens_per_game: 728188; score: 5%; 模型: GPT-5.6 Sol; 执行框架: Official harness; 推理强度: High; 得分: 5.2%; 每局输出词元数: 728,188; 输出词元总数: 18,204,692" role="graphics-symbol" aria-roledescription="point" transform="translate(100.90605142857143,233.16)" d="M4.183,0A4.183,4.183,0,1,1,-4.183,0A4.183,4.183,0,1,1,4.183,0" fill="currentColor" stroke-width="1.5" opacity="1"></path><path aria-label="output_tokens_per_game: 1285393; score: 7%; 模型: GPT-5.6 Sol; 执行框架: Official harness; 推理强度: Xhigh; 得分: 7.2%; 每局输出词元数: 1,285,393; 输出词元总数: 32,134,825" role="graphics-symbol" aria-roledescription="point" transform="translate(178.11874428571429,219.76000000000002)" d="M4.183,0A4.183,4.183,0,1,1,-4.183,0A4.183,4.183,0,1,1,4.183,0" fill="currentColor" stroke-width="1.5" opacity="1"></path><path aria-label="output_tokens_per_game: 2900997; score: 13%; 模型: GPT-5.6 Sol; 执行框架: Official harness; 推理强度: Max; 得分: 13.3%; 每局输出词元数: 2,900,997; 输出词元总数: 72,524,929" role="graphics-symbol" aria-roledescription="point" transform="translate(401.9952985714286,178.89)" d="M4.183,0A4.183,4.183,0,1,1,-4.183,0A4.183,4.183,0,1,1,4.183,0" fill="currentColor" stroke-width="1.5" opacity="1"></path></g></g></g></g><g role="graphics-object" aria-roledescription="text mark container"><text aria-label="output_tokens_per_game: 485485; score: 38%; endpoint_label: Harness with retained reasoning and compaction" role="graphics-symbol" aria-roledescription="text mark" text-anchor="start" transform="translate(79.27435,15.389999999999995)" font-family="OpenAI Sans, OpenAI Sans Variable Scripts, sans-serif" font-size="12px" fill="currentColor">Harness with retained reasoning and compaction</text> <text aria-label="output_tokens_per_game: 2900997; score: 13%; endpoint_label: Official harness" role="graphics-symbol" aria-roledescription="text mark" text-anchor="start" transform="translate(413.9952985714286,182.89)" font-family="OpenAI Sans, OpenAI Sans Variable Scripts, sans-serif" font-size="12px" fill="currentColor">Official harness</text></g></g></g></g></g></svg>

*使用官方执行框架时，GPT‑5.6 Sol 在 ARC-AGI-3 公开集上得分为 13.3%；保留推理并启用压缩后，得分达到 38.3%。分数衡量相对人类行动效率（* [*RHAE* ⁠](https://docs.arcprize.org/methodology) *）* — *该指标将模型表现与人类基准进行比较。根据* [*官方游戏记录* ⁠](https://huggingface.co/datasets/magic-sword/arc_agi_3_public_demo_human_testing) *，我们估计人类测试者的平均得分为 48%。模型不会获知评分方式，在整个过程中也看不到自己的分数* — *每次操作只会返回当前画面及所在关卡的文本表示。*

## ARC-AGI-3

ARC-AGI-3 是一项旨在衡量 AI 智能体学习和推理能力的基准测试。智能体需要探索陌生的二维游戏，在没有明确说明的情况下推断游戏机制。你可以前往 [arcprize.org/tasks⁠](https://arcprize.org/tasks) 体验 25 款演示游戏。

ARC-AGI-3 有意采用不含工具或特殊功能的通用执行框架。ARC 认为，简单的执行框架能更清楚地暴露模型的不足，也能让模型之间的比较更加公平。相比之下，商业开发者会针对每个模型的功能和特点优化执行框架。

在游戏方面，GPT‑5.6 Sol 已通过仅使用视觉的执行框架通关《宝可梦 火红》（由 [GPT\_Plays\_Pokemon⁠](https://www.twitch.tv/gpt_plays_pokemon) 直播），通过 Codex 的计算机操作功能通关《杀戮尖塔》（由 [EpochAI⁠](https://www.twitch.tv/epochaiplays) 直播），并完成《巴巴是你》的前几个关卡（由 [Piotr Migdał & Piotr Grabowski⁠](https://quesma.com/blog/baba-is-bench/) 分享）。ARC-AGI-3 究竟有何不同？

受到 [ARC 对 GPT‑5.5 缺陷的分析⁠](https://arcprize.org/blog/arc-agi-3-gpt-5-5-opus-4-7-analysis) 启发，我们查看了 GPT‑5.6 Sol 的一些尝试。和 ARC 一样，我们也发现模型看起来并不聪明。它每次行动前都会思考很久，却迟迟无法取得进展。

但在深入研究后，我们发现模型的困惑大多并非源于模型本身，而是执行框架中的设置。

首先，我们注意到，每次游戏操作后，所有私有推理都会被丢弃。这意味着 GPT‑5.6 Sol 每次行动时都必须重新理解游戏，无法记住之前的思考。模型仍能看到过去的操作记录和简短附注，但看不到促成这些操作的计划、见解或思考。

其次，我们发现执行框架采用了滚动截断窗口，随着历史记录增长，较早的操作会逐渐不可见。因此，GPT‑5.6 Sol 不仅无法记住过去的思考，也在逐渐忘记过去的操作。

执行框架的这两项特性——丢弃推理和滚动截断——共同解释了 GPT‑5.6 Sol 为何难以持续学习。

## 智能体记得自己做过什么时表现最佳

我们的模型经过训练，会先通过私有推理消息进行思考，再输出回复或调用工具。这些私有思考消息会作为对话历史的一部分保留下来。如果对话过长，我们会对其进行总结，然后继续。

我们的模型正是以这种方式训练，也以同样的方式部署在 ChatGPT 和 Codex 中。为了更贴近我们的生产环境，我们使用 [Responses API⁠](https://developers.openai.com/blog/responses-api) 实现了 ARC-AGI-3 执行框架。我们的 API 可轻松管理上下文：对于 GPT‑5.6，只需传入上一次响应的 ID，即可在工具调用和对话轮次之间自动保留推理。

保留推理后，我们观察到两项显著变化。首先，GPT‑5.6 Sol 每次行动前的思考时间更短，因为它不再需要每一轮都从头理解游戏。其次，能够记住过去的思考后，GPT‑5.6 Sol 更善于持续学习和采用连贯的策略。

ARC-AGI-3 执行框架通过滚动截断来应对上下文限制。当对话上下文超过 175,000 个字符时，最早的消息会被丢弃。

滚动截断有两个缺点。首先，模型会丢失较早的观察和操作。其次，模型在大部分任务期间都要使用接近满载的上下文窗口，这可能会略微影响表现。

在 ARC-AGI-3 上启用压缩后，GPT‑5.6 Sol 能在更长的运行过程中更好地保留从每款游戏中学到的内容，并以更少的输出词元取得更高的得分。

为了直观展示保留推理和启用压缩的效果，下面的动画呈现了 GPT‑5.6 Sol 分别使用两种执行框架解决一系列 ARC-AGI-3 谜题时，其 175K 上下文窗口的变化。

*中间两列展示了两种执行框架以不同方式使用模型上下文窗口的情况。由于能更好地记住过去，GPT‑5.6 Sol 每次行动所需的思考更少，推进速度也快得多。注意：我们的实现采用 175,000 个词元而非字符作为上限，但两者最终非常接近，因为绝大多数文本都是操作网格，我们的分词器会按 1:1 的比例将其转换为词元。*

保留推理与压缩结合使用后，GPT‑5.6 Sol（Max）的得分约为原来的三倍，而输出词元仅为原来的六分之一。

## 结论与建议

我们希望这些实验能够提醒大家：评测很少只衡量模型本身，也会衡量一系列不太显眼的选择，包括 API 设置、执行框架设计和提示方式。这并非我们第一次因公开基准测试中的低分而感到意外，随后才发现评测运行程序使用了会丢弃推理消息的通用执行框架。

如果你是希望充分发挥性能的 API 开发者，我们建议采用与自家产品相同的设置：

如果你要比较模型，我们建议采用使用上述设置的评测，因为这些设置最贴近 ChatGPT 和 Codex 中的实际应用。

我们感谢 ARC 多年来在 AGI 评测领域开展的创造性工作，也感谢其分析启发我们进一步深入研究。

如果你想与前沿模型一较高下，可以前往 [arcprize.org/tasks⁠](https://arcprize.org/tasks) 亲自挑战这些公开游戏。
