---
格式版本: 2
标题: "Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning"
原文链接: "https://arxiv.org/abs/2608.24658"
发布日期: "2026-08-25"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Tue, 25 Aug 2026 15:02:54 UTC (752 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-26T21:53:25+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-26T21:50:39+08:00"
入库时间: "2026-08-26T13:53:25.605Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=Scale-up&searchtype=all"
匹配关键词:
  - "Scale-up"
  - "latency"
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 37
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是通过子任务并行、试验并行和PA-GRPO加速LLM推理，属于模型推理效率研究，并非超节点或可独立复用的机架级通信、调度、网络、RAS机制。arXiv预印本为可追溯的一手学术来源但未经正式同行评审，当前页面仅提供摘要。相较给定历史证据，新增事实包括Trial Parallelism占DeepSeek-V4相关推理步骤65.5%及AIME基准约1.7倍加速，但没有机架架构、生产部署、客户或商业化信号；命中“应用与模型效率”及摘要壳否决项。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-08-26T21:53:36+08:00"
AI主题相关性: 2
AI来源权威性: 10
AI新颖性: 14
AI技术细节: 6
AI商业部署信号: 0
AI完整性: 5
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.3"
AI评分知识库SHA256: "dbc02c7552b478ae5aae514533e58a5de9ce72d2911e4aaac0f744c08178e4a6"
AI评分知识库检索词: "[\"Scale-up\",\"RAS\",\"Intel\",\"LLM\",\"PDF\",\"arxiv.org/pdf/2608.24658\",\"HTML\",\"arxiv.org/html/2608.24658v1\",\"LLMs\",\"DeepSeek-V4\",\"HLE\",\"PA-GRPO\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0131\",\"title\":\"DeepSeek-V4如何在昇腾超节点高效完成全参数后训练？SLAI T-Rex技术报告解读\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"LLM\",\"PDF\",\"HTML\",\"DeepSeek-V4\"],\"rank\":-12.165410380086012},{\"id\":\"july-correct-0081\",\"title\":\"AAI 2026: 6th Gen AMD EPYC Server CPUs Power the Agentic Data Center\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"RAS\",\"Intel\",\"PDF\",\"HTML\"],\"rank\":-10.190986881929163},{\"id\":\"july-correct-0111\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"LLM\",\"PDF\",\"HTML\"],\"rank\":-8.069899653483201},{\"id\":\"july-correct-0123\",\"title\":\"ODCC分享 | UALink联盟Kurtis：开放Scale-Up互连加速构建可部署AI超节点\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"Scale-up\",\"Intel\",\"LLM\",\"HTML\"],\"rank\":-7.818374723375065},{\"id\":\"july-correct-0130\",\"title\":\"ODCC分享 | UALink联盟Kurtis：开放Scale-Up互连加速构建可部署AI超节点\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"Scale-up\",\"Intel\",\"LLM\",\"HTML\"],\"rank\":-7.818374723375065}]"
AI摘要: "Parason 提出一种并行推理框架，将 LLM 长推理中的子任务并行与试验并行统一建模，并发现试验并行在 DeepSeek-V4 推理中占可并行计算的 65.5%。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-08-27T08:12:46.794Z"
采集批次: "2026年8月26日20点37分07秒"
采集批次ID: "20260826-203707-934"
去重键: "https://arxiv.org/abs/2608.24658"
---

## Computer Science > Artificial Intelligence

## Title:Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning

Authors:[Zhengyang Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+Z), [Zijian Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+Z), [Jiaxuan Gao](https://arxiv.org/search/cs?searchtype=author&query=Gao,+J), [Shusheng Xu](https://arxiv.org/search/cs?searchtype=author&query=Xu,+S), [Yi Wu](https://arxiv.org/search/cs?searchtype=author&query=Wu,+Y), [Song Han](https://arxiv.org/search/cs?searchtype=author&query=Han,+S), [Ligeng Zhu](https://arxiv.org/search/cs?searchtype=author&query=Zhu,+L)

[View PDF](https://arxiv.org/pdf/2608.24658) [HTML (experimental)](https://arxiv.org/html/2608.24658v1)

> Abstract:Scaling test-time reasoning has substantially improved the problem-solving ability of large language models (LLMs), but standard autoregressive decoding still executes long reasoning traces sequentially, creating severe latency for difficult tasks (up to days and weeks). Parallel reasoning offers a natural remedy. However, prior systems primarily focus on Subtask Parallelism, where the model learns to decompose a high-level task into smaller chunks that can be solved independently. This approach overlooks another pervasive form of parallelism: Trial Parallelism, where multiple speculative attempts explore, verify, and aggregate competing hypotheses in parallel. In this paper, we introduce Parason, which reveals and learns both forms of parallelism in LLM reasoning. Our analysis identifies Trial Parallelism as the majority of parallelizable reasoning computation (65.5% in DeepSeek-V4's reasoning steps in HLE), and it becomes increasingly dominant on hard problems. Guided by this taxonomy, Parason converts sequential reasoning traces into structured parallel trajectories with a context-free grammar, then trains models with Parallelism-Aware Group Relative Policy Optimization (PA-GRPO), whose reward jointly balances accuracy, latency, and the two parallelism ratios. At inference time, Parason executes the learned parallel structure through tool calls, translating theoretical savings to real-world wall-clock acceleration. Experiments on mathematical reasoning benchmarks including AIME24 and AIME25 show that Parason achieves an average acceleration about 1.7 $\times$ while maintaining competitive accuracy.

| Subjects: | Artificial Intelligence (cs.AI) |
| --- | --- |
| Cite as: | [arXiv:2608.24658](https://arxiv.org/abs/2608.24658) \[cs.AI\] |
|  | (or [arXiv:2608.24658v1](https://arxiv.org/abs/2608.24658v1) \[cs.AI\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.24658](https://doi.org/10.48550/arXiv.2608.24658) |

## Submission history

From: Zhengyang Zhang \[[view email](https://arxiv.org/show-email/179e7224/2608.24658)\]  
**\[v1\]** Tue, 25 Aug 2026 15:02:54 UTC (752 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.24658) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
