---
格式版本: 2
标题: "WISE: World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models"
原文链接: "https://arxiv.org/abs/2609.03681"
发布日期: "2026-09-03"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Thu, 3 Sep 2026 11:17:57 UTC (18,313 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-06T21:24:06+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-06T21:21:19+08:00"
入库时间: "2026-09-06T13:24:06.111Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=GPU&searchtype=all"
匹配关键词:
  - "GPU"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 7
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "该论文为机器人视觉-语言-动作模型训练研究，仅提及GPU计算时间，与超节点、AI Rack、机柜级基础设施、供电散热互连等核心主题完全无关。"
AI质检模型: "zj-deepseek-v4-flash"
AI质检时间: "2026-09-06T21:24:11+08:00"
AI主题相关性: 0
AI来源权威性: 2
AI新颖性: 0
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 5
AI摘要: "这篇论文提出WISE框架，通过世界模型引导的想象调度，在交互关键状态选择性调用有界多视图想象，并以进展和完成信号评估候选未来，用于视觉-语言-动作模型的高效后训练。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-06T23:44:23.520Z"
采集批次: "2026年9月6日19点27分09秒"
采集批次ID: "20260906-192709-237"
去重键: "https://arxiv.org/abs/2609.03681"
---

## Computer Science > Robotics

## Title:WISE: World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models

Authors:[Chenhao Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+C), [Hanyu Zhao](https://arxiv.org/search/cs?searchtype=author&query=Zhao,+H), [Hang Cheng](https://arxiv.org/search/cs?searchtype=author&query=Cheng,+H), [Tengfei Pan](https://arxiv.org/search/cs?searchtype=author&query=Pan,+T), [Long Zeng](https://arxiv.org/search/cs?searchtype=author&query=Zeng,+L)

[View PDF](https://arxiv.org/pdf/2609.03681) [HTML (experimental)](https://arxiv.org/html/2609.03681v1)

> Abstract:Post-training VLA policies typically rely on supervised fine-tuning with costly expert demonstrations or reinforcement learning with expensive and potentially unstable real-world exploration. World models offer a promising alternative by evaluating candidate behaviors through imagined futures, yet effective post-training requires more than accurate prediction: imagination must be scheduled where it is useful, bounded within reliable horizons, and translated into trustworthy policy supervision. In robotic manipulation, the value of imagination varies substantially across execution stages, while extended rollouts can accumulate prediction errors and introduce unreliable learning signals. We introduce WISE (World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models), a unified framework that coordinates when and how world-model imagination is used during policy refinement. WISE selectively invokes imagination at interaction-relevant states, performs bounded multi-view rollouts, evaluates candidate futures using progress and completion signals, and uses their relative outcomes to refine actions generated from real interaction contexts. Extensive experiments with both $\pi_0$ and $\pi_{0.5}$ demonstrate consistent improvements across diverse manipulation tasks while reducing GPU computation time by approximately 80% compared with full imagination. Real-world evaluations further show substantial gains in robustness and generalization under diverse real-world distribution shifts.

| Subjects: | Robotics (cs.RO) |
| --- | --- |
| Cite as: | [arXiv:2609.03681](https://arxiv.org/abs/2609.03681) \[cs.RO\] |
|  | (or [arXiv:2609.03681v1](https://arxiv.org/abs/2609.03681v1) \[cs.RO\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.03681](https://doi.org/10.48550/arXiv.2609.03681) |

## Submission history

From: Chenhao Zhang \[[view email](https://arxiv.org/show-email/24c5f644/2609.03681)\]  
**\[v1\]** Thu, 3 Sep 2026 11:17:57 UTC (18,313 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.03681) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
