---
格式版本: 2
标题: "Partially Observable Learning for Multi-Platform Dispatch Optimization"
原文链接: "https://arxiv.org/abs/2608.10897"
发布日期: "2026-08-11"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Tue, 11 Aug 2026 13:20:12 UTC (7,467 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-12T11:14:08+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-12T11:12:44+08:00"
入库时间: "2026-08-12T03:14:08.327Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=delivery&searchtype=all"
匹配关键词:
  - "delivery"
  - "performance"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 5
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "论文讨论即时配送平台调度优化，与超节点/AI Rack/机柜级AI基础设施完全无关，核心主题不匹配。"
AI质检模型: "ali-deepseek-v4-flash"
AI质检时间: "2026-08-12T11:14:19+08:00"
AI主题相关性: 0
AI来源权威性: 5
AI新颖性: 0
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 0
采集批次: "2026年8月12日2点04分20秒"
采集批次ID: "20260812-020420-289"
去重键: "https://arxiv.org/abs/2608.10897"
---

## Computer Science > Machine Learning

## Title:Partially Observable Learning for Multi-Platform Dispatch Optimization

Authors:[Fengming Yao](https://arxiv.org/search/cs?searchtype=author&query=Yao,+F), [Man Luo](https://arxiv.org/search/cs?searchtype=author&query=Luo,+M)

[View PDF](https://arxiv.org/pdf/2608.10897) [HTML (experimental)](https://arxiv.org/html/2608.10897v1)

> Abstract:Instant delivery platforms have become a critical component of urban logistics, increasingly relying on crowdsourced couriers to fulfill highly dynamic orders. In real-world systems, couriers are not exclusive to a single platform and may concurrently serve multiple platforms, while each platform can only observe its own orders and couriers' interactions due to privacy and operational constraints. This results in a multi-platform dispatch environment with inherent partial observability. However, most existing works on dispatch optimization assume full courier observability and mandatory assignment acceptance, causing substantial performance degradation when deployed in realistic multi-platform settings. In this paper, we propose POLO, a partially observable multi-agent reinforcement learning framework for dispatching optimization in multi-platform instant delivery systems. POLO firstly models each platform-grid pair as an independent agent that learns dispatch policies solely from platform-local observations, aligning the learning process with real-world privacy and operational constraints. To support effective decision-making under incomplete and heterogeneous courier information, POLO introduces a novel attention-based policy representation that selectively aggregates inter-courier information. Moreover, we design a counterfactual reward shaping mechanism to mitigate the non-stationarity induced by joint actions across grids, leading to more stable and scalable learning. We develop a high-fidelity simulator to evaluate dispatch performance under varying numbers of platforms and system scales. Extensive experiments demonstrate that POLO consistently outperforms strong baselines in terms of platform revenue and courier travel efficiency, highlighting its robustness and effectiveness in realistic multi-platform settings.

| Comments: |  |
| --- | --- |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | [arXiv:2608.10897](https://arxiv.org/abs/2608.10897) \[cs.LG\] |
|  | (or [arXiv:2608.10897v1](https://arxiv.org/abs/2608.10897v1) \[cs.LG\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.10897](https://doi.org/10.48550/arXiv.2608.10897) |

## Submission history

From: Fengming Yao \[[view email](https://arxiv.org/show-email/482d529a/2608.10897)\]  
**\[v1\]** Tue, 11 Aug 2026 13:20:12 UTC (7,467 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.10897) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
