---
格式版本: 2
标题: "Boosting LLM Exploration via Weak-Model Guidance in RLVR"
原文链接: "https://arxiv.org/abs/2608.27420"
发布日期: "2026-08-27"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Thu, 27 Aug 2026 17:45:50 UTC (291 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-30T15:12:13+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-30T15:11:38+08:00"
入库时间: "2026-08-30T07:12:13.333Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=Scale-up&searchtype=all"
匹配关键词:
  - "Scale-up"
  - "performance"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 32
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是通过弱模型生成的部分推理轨迹改善RLVR探索与熵坍缩，属于模型训练算法研究，不涉及超节点、AI机架、互连、供电、散热或机架级部署。来源为作者提交的arXiv原始预印本摘要，尚非正式标准或生产部署材料。固定知识库未显示同一方法，但未命中不能证明首次出现；本文可识别的新增仅是弱模型前缀引导方法及数学基准实验结论。无客户、量产、交付或生产平台信号，且命中“应用与模型效率”硬否决项，当前页面不值得作为超节点业务信息源。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-08-30T15:12:22+08:00"
AI主题相关性: 0
AI来源权威性: 9
AI新颖性: 11
AI技术细节: 5
AI商业部署信号: 0
AI完整性: 7
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.55"
AI评分知识库SHA256: "ca56afec1ff616b1fa4394cd1d08b1e2a6af47eccbe70c802c9f1fea6037e803"
AI评分知识库检索词: "[\"Scale-up\",\"LLM\",\"RLVR\",\"PDF\",\"arxiv.org/pdf/2608.27420\",\"HTML\",\"arxiv.org/html/2608.27420v1\",\"LLMs\",\"SFT\",\"arxiv.org/abs/2608.27420\",\"arxiv.org/abs/2608.27420v1\",\"doi.org/10.48550/arXiv.2608.27420\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0131\",\"title\":\"DeepSeek-V4如何在昇腾超节点高效完成全参数后训练？SLAI T-Rex技术报告解读\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"LLM\",\"PDF\",\"HTML\",\"SFT\"],\"rank\":-11.776979353178586},{\"id\":\"runtime-569639751a0dbe7ef3ffccc5\",\"title\":\"[2608.17503] Predict Before Replay: Joint FEC and Flight Control for Reliable Scale-Up Links\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-18\",\"matchedTerms\":[\"Scale-up\",\"PDF\",\"HTML\"],\"rank\":-10.481970259177816},{\"id\":\"july-correct-0111\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"LLM\",\"PDF\",\"HTML\"],\"rank\":-8.840966575116735},{\"id\":\"historical-jun-010\",\"title\":\"爱建证券-电子行业专题报告：Vera Rubin量产提速，RTX Spark打开终端AI新空间-260608.pdf\",\"sourceType\":\"curated_item\",\"time\":\"2026-06\",\"matchedTerms\":[\"PDF\"],\"rank\":-6.203557389986544},{\"id\":\"runtime-5042ddb06fd844309130b6e9\",\"title\":\"Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-12\",\"matchedTerms\":[\"LLM\",\"SFT\"],\"rank\":-5.929470497032034}]"
AI摘要: "该文提出在RLVR训练中引入较弱小模型生成的部分推理轨迹作为外部前缀，引导目标LLM探索不同推理路径，以缓解策略熵下降和推理覆盖收窄问题。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-08-30T18:08:58.889Z"
采集批次: "2026年8月30日14点11分37秒"
采集批次ID: "20260830-141137-130"
去重键: "https://arxiv.org/abs/2608.27420"
---

## Computer Science > Computation and Language

## Title:Boosting LLM Exploration via Weak-Model Guidance in RLVR

Authors:[Xingyu Shen](https://arxiv.org/search/cs?searchtype=author&query=Shen,+X), [Huishuai Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+H), [Peng Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+P), [Yinchun Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+Y), [Dongyan Zhao](https://arxiv.org/search/cs?searchtype=author&query=Zhao,+D)

[View PDF](https://arxiv.org/pdf/2608.27420) [HTML (experimental)](https://arxiv.org/html/2608.27420v1)

> Abstract:Reinforcement Learning with Verifiable Rewards (RLVR) significantly improves LLM reasoning but often causes a drop in policy entropy, leading to narrowed reasoning coverage and degraded pass@$k$ for large $k$. While existing methods mitigate this entropy collapse through algorithmic regularizations, cross-model non-parametric perturbation is also neglected. In this work, we propose a simple yet effective approach to preserve the generative diversity of LLMs during RLVR. Instead of relying solely on internal exploration, we force the target model to generate answers based on partial reasoning trajectories generated by a smaller, weaker language models. These unfamiliar prefixes effectively disrupt over-confidence and encourage the exploration of distinct reasoning paths. We empirically study the potential of outer prefixes, revealing the mechanism of the impact of distributional discrepancy to the exploration dynamics in RLVR training. Experiments across multiple mathematical benchmarks show that our method consistently outperforms vanilla RLVR. Notably, the performance gain becomes increasingly pronounced as $k$ scales up, demonstrating a substantial expansion of reasoning coverage. Furthermore, our approach efficiently mitigates entropy collapse without requiring additional SFT, intricate reward designs, or complex prompting.

| Comments: |  |
| --- | --- |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | [arXiv:2608.27420](https://arxiv.org/abs/2608.27420) \[cs.CL\] |
|  | (or [arXiv:2608.27420v1](https://arxiv.org/abs/2608.27420v1) \[cs.CL\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.27420](https://doi.org/10.48550/arXiv.2608.27420) |

## Submission history

From: Xingyu Shen \[[view email](https://arxiv.org/show-email/837c92b9/2608.27420)\]  
**\[v1\]** Thu, 27 Aug 2026 17:45:50 UTC (291 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.27420) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
