---
格式版本: 2
标题: "Simple Actors and Deep Critics for Scalable Reinforcement Learning"
原文链接: "https://arxiv.org/abs/2608.26659"
发布日期: "2026-08-27"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Thu, 27 Aug 2026 06:14:12 UTC (1,293 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-30T15:14:04+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-30T15:11:38+08:00"
入库时间: "2026-08-30T07:14:04.179Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=Scale-up&searchtype=all"
匹配关键词:
  - "Scale-up"
  - "deployment"
  - "latency"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 28
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是离线强化学习算法LAC，通过轻量actor和深层critic降低推理延迟，与超节点、AI机架及机架级基础设施无关。来源为arXiv一手预印本摘要，新增残差MLP、n-step目标、分类交叉熵及最高4倍延迟改善等研究结果；历史证据未见同一论文，但这些新增仅属于模型算法。无机架架构、互连、供电、散热、生产部署或商业信号，命中“应用与模型效率”强否决项。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-08-30T15:14:37+08:00"
AI主题相关性: 0
AI来源权威性: 10
AI新颖性: 9
AI技术细节: 2
AI商业部署信号: 0
AI完整性: 7
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.55"
AI评分知识库SHA256: "ca56afec1ff616b1fa4394cd1d08b1e2a6af47eccbe70c802c9f1fea6037e803"
AI评分知识库检索词: "[\"Scale-up\",\"PDF\",\"arxiv.org/pdf/2608.26659\",\"HTML\",\"arxiv.org/html/2608.26659v1\",\"v1\",\"RL\",\"MLP\",\"LAC\",\"OGBench\",\"arxiv.org/abs/2608.26659\",\"arxiv.org/abs/2608.26659v1\"]"
AI评分知识库命中: "[{\"id\":\"runtime-569639751a0dbe7ef3ffccc5\",\"title\":\"[2608.17503] Predict Before Replay: Joint FEC and Flight Control for Reliable Scale-Up Links\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-18\",\"matchedTerms\":[\"Scale-up\",\"PDF\",\"HTML\",\"v1\",\"RL\"],\"rank\":-12.808968674720035},{\"id\":\"july-correct-0111\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"PDF\",\"HTML\",\"v1\",\"RL\"],\"rank\":-11.333006229301258},{\"id\":\"july-correct-0010\",\"title\":\"Schneider Electric and AMD release first Helios platform reference design to accelerate AI Factory deployment\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"PDF\",\"v1\",\"RL\"],\"rank\":-8.011504167715488},{\"id\":\"july-correct-0094\",\"title\":\"NVIDIA Vera CPU: Olympus Cores Built for Maximum Single-Thread Performance in Agentic AI\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"Scale-up\",\"RL\",\"LAC\"],\"rank\":-7.328187324161346},{\"id\":\"july-correct-0115\",\"title\":\"锚定 300kW 整机柜演进方向 OAII 社区三项规范联合发布，树立AIDC基础设施建设标准答案\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"HTML\",\"v1\"],\"rank\":-6.992347863111062}]"
AI摘要: "离线强化学习论文提出 LAC（轻量 Actor、深度 Critic）方法，针对扩散/流匹配策略等生成式 Actor 部署开销大的问题，将模型容量更多分配给仅在训练时使用的深度 Critic，并分别解决深层 Critic 的优化。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-08-30T18:08:40.652Z"
采集批次: "2026年8月30日14点11分37秒"
采集批次ID: "20260830-141137-130"
去重键: "https://arxiv.org/abs/2608.26659"
---

## Computer Science > Machine Learning

## Title:Simple Actors and Deep Critics for Scalable Reinforcement Learning

Authors:[Guhyeon Kang](https://arxiv.org/search/cs?searchtype=author&query=Kang,+G), [Jaehwi Lee](https://arxiv.org/search/cs?searchtype=author&query=Lee,+J), [Minhae Kwon](https://arxiv.org/search/cs?searchtype=author&query=Kwon,+M)

[View PDF](https://arxiv.org/pdf/2608.26659) [HTML (experimental)](https://arxiv.org/html/2608.26659v1)

> Abstract:Recent progress in offline reinforcement learning (RL) has been driven by expressive generative actors such as diffusion and flow-matching policies, which capture multimodal behavior in offline datasets. However, these actors require multiple denoising or integration steps per action and thus incur substantial overhead at every decision in deployment. In this work, we revisit where capacity should be invested in an offline actor--critic method. Since the critic is used only during training and is discarded at deployment while the actor runs at every decision step, allocating capacity to the critic rather than the actor is more favorable for inference-time efficiency. However, scaling MLP critics in offline RL is known to introduce several distinct instabilities that have, in practice, kept critics shallow. We identify three distinct failure modes that arise when critics are deepened in offline RL---optimization, bootstrap-noise amplification, and value-range drift---and address each with a corresponding ingredient: a residual MLP backbone, n-step bootstrap targets, and a categorical cross-entropy loss. Combining these ingredients with a lightweight deterministic actor, we propose LAC (Light Actor, deep Critic). On OGBench, LAC matches the strongest diffusion- and flow-matching baselines while achieving up to 4x lower inference latency, comparable to one-step distilled policies without distillation. Its critic recipe also transfers across actor parametrizations.

| Comments: |  |
| --- | --- |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | [arXiv:2608.26659](https://arxiv.org/abs/2608.26659) \[cs.LG\] |
|  | (or [arXiv:2608.26659v1](https://arxiv.org/abs/2608.26659v1) \[cs.LG\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.26659](https://doi.org/10.48550/arXiv.2608.26659) |

## Submission history

From: Guhyeon Kang \[[view email](https://arxiv.org/show-email/df0ae98e/2608.26659)\]  
**\[v1\]** Thu, 27 Aug 2026 06:14:12 UTC (1,293 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.26659) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
