---
格式版本: 2
标题: "FlashVLA: Streaming Action Decoding for Fast and Asynchronous VLA Inference"
原文链接: "https://arxiv.org/abs/2608.27384"
发布日期: "2026-08-27"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Thu, 27 Aug 2026 17:19:29 UTC (3,873 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-30T15:09:50+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-30T15:08:58+08:00"
入库时间: "2026-08-30T07:09:50.741Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=GPU&searchtype=all"
匹配关键词:
  - "GPU"
  - "deployment"
  - "performance"
  - "latency"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 36
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是机器人VLA模型的流式动作解码与单GPU推理优化，并非超节点或机架级AI基础设施。来源为作者提交的arXiv预印本摘要页，属于原始研究但未经同行评审且缺少论文全文细节。固定知识库未发现FlashVLA同项历史记录，但未命中不能证明首次出现；本文可核验新增为分块因果注意力、流式动作缓冲及单GPU不低于30Hz控制频率。没有客户、量产或机架级部署信号，命中“应用与模型效率”强否决项，不能进入业务情报库。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-08-30T15:10:13+08:00"
AI主题相关性: 1
AI来源权威性: 10
AI新颖性: 13
AI技术细节: 5
AI商业部署信号: 1
AI完整性: 6
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.55"
AI评分知识库SHA256: "ca56afec1ff616b1fa4394cd1d08b1e2a6af47eccbe70c802c9f1fea6037e803"
AI评分知识库检索词: "[\"GPU\",\"VLA\",\"PDF\",\"arxiv.org/pdf/2608.27384\",\"HTML\",\"arxiv.org/html/2608.27384v1\",\"v1\",\"VLM\",\"arxiv.org/abs/2608.27384\",\"arxiv.org/abs/2608.27384v1\",\"doi.org/10.48550/arXiv.2608.27384\",\"arxiv.org/show-email/3018df9c/2608.27384\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0111\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"PDF\",\"HTML\",\"v1\"],\"rank\":-11.333006229301258},{\"id\":\"runtime-569639751a0dbe7ef3ffccc5\",\"title\":\"[2608.17503] Predict Before Replay: Joint FEC and Flight Control for Reliable Scale-Up Links\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-18\",\"matchedTerms\":[\"VLA\",\"PDF\",\"HTML\",\"v1\"],\"rank\":-10.52985569710448},{\"id\":\"july-correct-0010\",\"title\":\"Schneider Electric and AMD release first Helios platform reference design to accelerate AI Factory deployment\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"GPU\",\"PDF\",\"v1\"],\"rank\":-8.011504167715488},{\"id\":\"july-correct-0115\",\"title\":\"锚定 300kW 整机柜演进方向 OAII 社区三项规范联合发布，树立AIDC基础设施建设标准答案\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"HTML\",\"v1\"],\"rank\":-6.992347863111062},{\"id\":\"historical-may-043\",\"title\":\"2026安全可靠测评结果(第2号)出炉，壁仞科技壁砺166获I级认证\",\"sourceType\":\"curated_item\",\"time\":\"2026-05\",\"matchedTerms\":[\"GPU\",\"v1\"],\"rank\":-6.432979617351396}]"
AI摘要: "FlashVLA提出一种流式动作解码框架，面向流匹配类视觉-语言-动作（VLA）模型的推理高延迟与异步执行不稳定问题，通过维护多噪声级别的流式动作缓冲并采用分块因果注意力，使每个推理步骤产出一个可执行动作块并保持动作连续性。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-08-30T18:09:02.231Z"
采集批次: "2026年8月30日14点11分37秒"
采集批次ID: "20260830-141137-130"
去重键: "https://arxiv.org/abs/2608.27384"
---

## Computer Science > Robotics

## Title:FlashVLA: Streaming Action Decoding for Fast and Asynchronous VLA Inference

Authors:[Zekai Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+Z), [Jiaming Tang](https://arxiv.org/search/cs?searchtype=author&query=Tang,+J), [Zhijian Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+Z)

[View PDF](https://arxiv.org/pdf/2608.27384) [HTML (experimental)](https://arxiv.org/html/2608.27384v1)

> Abstract:Vision-Language-Action (VLA) models are increasingly promising for robotic manipulation, yet their real-world deployment remains bottlenecked by high inference latency and unstable asynchronous execution. This challenge is particularly pronounced in flow-matching-based VLA models, where action decoding requires multiple iterative steps conditioned on the VLM context. While efficient inference methods improve control frequency and asynchronous methods reduce execution idle time, existing approaches often fail to jointly achieve low-latency inference and accurate, temporally consistent asynchronous execution. We introduce \\textbf{FlashVLA}, a streaming action decoding framework that addresses both challenges in a unified formulation. FlashVLA maintains a streaming action buffer with multiple chunks at different noise levels and decodes them using chunk-wise causal attention. This design allows FlashVLA to produce one executable action chunk per inference step. Moreover, its chunk-wise autoregressive formulation implicitly preserves action continuity, enabling smooth asynchronous execution without extra future-state conditioning. Across extensive simulated and real-world experiments, FlashVLA substantially improves inference speed while maintaining strong task performance. It can achieve $\geq$ 30\\,Hz control frequency on a single GPU with smooth asynchronous inference in real-world deployment.

| Comments: |  |
| --- | --- |
| Subjects: | Robotics (cs.RO) |
| Cite as: | [arXiv:2608.27384](https://arxiv.org/abs/2608.27384) \[cs.RO\] |
|  | (or [arXiv:2608.27384v1](https://arxiv.org/abs/2608.27384v1) \[cs.RO\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.27384](https://doi.org/10.48550/arXiv.2608.27384) |

## Submission history

From: ZeKai Li \[[view email](https://arxiv.org/show-email/3018df9c/2608.27384)\]  
**\[v1\]** Thu, 27 Aug 2026 17:19:29 UTC (3,873 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.27384) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
