---
格式版本: 2
标题: "Pallas: A Proactive KV Cache Migration Framework for LLM Inference in AI-RAN"
原文链接: "https://arxiv.org/abs/2608.16477"
发布日期: "2026-08-17"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Mon, 17 Aug 2026 12:16:09 UTC (439 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-20T15:19:02+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-20T15:12:40+08:00"
入库时间: "2026-08-20T07:19:02.317Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=delivery&searchtype=all"
匹配关键词:
  - "delivery"
  - "latency"
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 26
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "论文主题为AI-RAN场景下LLM推理的KV Cache迁移，属于边缘推理与移动网络优化，未涉及超节点、机柜级AI基础设施、硬件架构、供电散热或量产落地等核心方向，与本项目主题无关。来源为arXiv学术预印本，权威性一般，无商业部署信号。"
AI质检模型: "ali-deepseek-v4-flash"
AI质检时间: "2026-08-20T15:20:51+08:00"
AI主题相关性: 2
AI来源权威性: 8
AI新颖性: 5
AI技术细节: 3
AI商业部署信号: 0
AI完整性: 8
AI摘要: "Pallas提出一种面向AI-RAN的主动式KV缓存迁移框架，在移动用户切换基站前预测目标并并行准备推理状态：目标基站通过本地prefill重建历史前缀，源端流式传输动态后缀的KV块。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T03:25:08.949Z"
采集批次: "2026年8月20日14点19分32秒"
采集批次ID: "20260820-141932-079"
去重键: "https://arxiv.org/abs/2608.16477"
---

## Computer Science > Machine Learning

## Title:Pallas: A Proactive KV Cache Migration Framework for LLM Inference in AI-RAN

Authors:[Tianhang Ding](https://arxiv.org/search/cs?searchtype=author&query=Ding,+T), [Jianchun Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+J), [Hongli Xu](https://arxiv.org/search/cs?searchtype=author&query=Xu,+H)

[View PDF](https://arxiv.org/pdf/2608.16477) [HTML (experimental)](https://arxiv.org/html/2608.16477v1)

> Abstract:AI-RAN brings large language model (LLM) serving close to mobile users, but cellular handover can separate an active request from its inference state: the user attaches to a target base station (gNB) while the large and growing key-value (KV) cache remains at the source. Retaining inference at the source preserves service continuity but persistently increases inter-token latency (ITL), whereas recovering the state at the target restores serving locality but requires KV-cache transfer, recomputation, or a combination of both only after handover, directly prolonging service interruption time (SIT).  
> This work presents Pallas, a \\textit{proactive} KV-cache migration framework that prepares the inference state at the predicted target before handover, in parallel with ongoing source-side inference and token delivery. At the preparation trigger, Pallas partitions the token sequence into a stable historical prefix and an evolving suffix. The target reconstructs the prefix through local prefill, while the source streams the KV blocks generated for the suffix. At handover, the target assembles both portions into an up-to-date KV cache and resumes decoding locally, leaving only unfinished preparation to contribute to SIT. An online scheduler selects the \\textit{prefetching window}, which determines how early preparation begins before handover, based on mobility predictions and runtime telemetry. Across three LLMs and $100$-- $500~\mathrm{Mbps}$ inter-gNB links, our vLLM-based prototype reduces average SIT by factors of $2.28$--$89.68$ over target-side recovery approaches and lowers average ITL by $16.0\\%$--$50.0\\%$ compared with source-side forwarding.

| Subjects: | Machine Learning (cs.LG) |
| --- | --- |
| Cite as: | [arXiv:2608.16477](https://arxiv.org/abs/2608.16477) \[cs.LG\] |
|  | (or [arXiv:2608.16477v1](https://arxiv.org/abs/2608.16477v1) \[cs.LG\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.16477](https://doi.org/10.48550/arXiv.2608.16477) |

## Submission history

From: Jianchun Liu \[[view email](https://arxiv.org/show-email/9d4358bc/2608.16477)\]  
**\[v1\]** Mon, 17 Aug 2026 12:16:09 UTC (439 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.16477) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
