---
格式版本: 2
标题: "HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting"
原文链接: "https://arxiv.org/abs/2608.11692"
发布日期: "2026-08-12"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 12 Aug 2026 06:04:27 UTC (2,942 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-13T18:34:23+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-13T18:24:32+08:00"
入库时间: "2026-08-13T10:34:23.449Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=AI&searchtype=all"
匹配关键词:
  - "AI"
  - "deployment"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI质检状态: "评分失败"
AI评分尝试次数: 1
AI评分错误类型: "service_error"
AI评分错误: "LLM call failed; tried model chain: ali-deepseek-v4-flash -> tx-deepseek-v4-flash | Model ali-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131034257834135058268d9d6bx325tKs)\",\"type\":\"new_api_error\"}} | Model tx-deepseek-v4-flash failed 503: {\"error\":{\"code\":\"model_not_found\",\"message\":\"No available channel for model tx-deepseek-v4-flash under group vip (distributor) (request id: 202608131034276154774018268d9d6VU2"
AI评分开始时间: "2026-08-13T10:34:23.460Z"
AI评分结束时间: "2026-08-13T10:34:27.739Z"
AI摘要: "本文提出 HUGIN 训练框架，用于提升视觉语言模型在自主物流分拣中的多场景联合规划能力；"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T03:39:59.402Z"
采集批次: "2026年8月13日18点24分27秒"
采集批次ID: "20260813-182427-1564e572"
去重键: "https://arxiv.org/abs/2608.11692"
---

## Computer Science > Artificial Intelligence

## Title:HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting

Authors:[Xikai Sun](https://arxiv.org/search/cs?searchtype=author&query=Sun,+X), [Cangtian Zhou](https://arxiv.org/search/cs?searchtype=author&query=Zhou,+C), [Kebin Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+K), [Ke Ma](https://arxiv.org/search/cs?searchtype=author&query=Ma,+K), [Xu Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+X), [Zaishu Chen](https://arxiv.org/search/cs?searchtype=author&query=Chen,+Z), [Haotian Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+H), [Li Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+L), [Yunhao Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+Y)

[View PDF](https://arxiv.org/pdf/2608.11692) [HTML (experimental)](https://arxiv.org/html/2608.11692v1)

> Abstract:Autonomous logistics sorting systems (ALSS) are an important industrial application of embodied AI, which requires joint planning over spatially disjoint camera views. We formulate this setting as Joint Multi-Scene Understanding (JMSU). With open-world visual understanding and task-planning capabilities, vision-language models (VLMs) are promising candidates for JMSU. However, directly applying existing VLMs to JMSU is non-trivial due to scarce cross-scene supervision and attention dispersion caused by long visual context in JMSU. To address these challenges, we propose HUGIN, a training framework with two complementary components. Endogenous Data Augmentation recombines verified atomic facts under operating constraints, while Global Context Ranking aligns the instruction representation more strongly with the complete visual context than with a partial visual context. To support ongoing research, we construct a high-quality industrial sorting dataset and benchmark named SortingBench from four layouts of autonomous logistics sorting systems. Across five open VLMs, HUGIN consistently outperforms matched baselines; for example, the accuracy on SortingBench of Qwen3-VL-8B increases from 63.6% to 78.8%. Additional experiments verify the effectiveness of each component and JMSU's spillover benefits in embodied tasks. Deployment tests involving more than 15,000 packages support the practical viability of VLM-based planning for autonomous logistics sorting.

| Subjects: | Artificial Intelligence (cs.AI) |
| --- | --- |
| Cite as: | [arXiv:2608.11692](https://arxiv.org/abs/2608.11692) \[cs.AI\] |
|  | (or [arXiv:2608.11692v1](https://arxiv.org/abs/2608.11692v1) \[cs.AI\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.11692](https://doi.org/10.48550/arXiv.2608.11692) |

## Submission history

From: Xikai Sun \[[view email](https://arxiv.org/show-email/6a706ac0/2608.11692)\]  
**\[v1\]** Wed, 12 Aug 2026 06:04:27 UTC (2,942 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.11692) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
