---
格式版本: 2
标题: "CacheBridge: Efficient Cross-Model KV Cache Transfer"
原文链接: "https://arxiv.org/abs/2609.00891"
发布日期: "2026-09-01"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Tue, 1 Sep 2026 08:20:35 UTC (239 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-02T18:52:03+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-02T18:51:36+08:00"
入库时间: "2026-09-02T10:52:03.276Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=GPU&searchtype=all"
匹配关键词:
  - "GPU"
  - "deployment"
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 49
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是跨模型KV Cache迁移算法，新增按架构匹配注意力头、注意力敏感校准和融合GPU核，并报告8倍存储缩减、最高3倍应用加速及10.7倍构建加速；arXiv预印本具备基本可追溯性，但当前页面仅有摘要。固定知识库未显示同项成果，不过未命中不能证明首次出现。该工作主要优化模型推理与缓存转换，未提供可独立复用的机架级通信、调度、RAS或网络机制，也无生产平台、客户、部署或商业化证据，命中“应用与模型效率”及摘要壳否决项。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-09-02T18:52:38+08:00"
AI主题相关性: 3
AI来源权威性: 11
AI新颖性: 15
AI技术细节: 15
AI商业部署信号: 0
AI完整性: 5
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.81"
AI评分知识库SHA256: "3b93d12e47749b3512f545f51c44c011bdc0931677c2cfe61e4df59a2b1a5a48"
AI评分知识库检索词: "[\"GPU\",\"Intel\",\"KV\",\"PDF\",\"arxiv.org/pdf/2609.00891\",\"HTML\",\"arxiv.org/html/2609.00891v1\",\"LLMs\",\"Qwen3\",\"to32\",\"arxiv.org/abs/2609.00891\",\"arxiv.org/abs/2609.00891v1\"]"
AI评分知识库命中: "[{\"id\":\"runtime-5042ddb06fd844309130b6e9\",\"title\":\"Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-12\",\"matchedTerms\":[\"GPU\",\"KV\",\"Qwen3\"],\"rank\":-12.37934374409118},{\"id\":\"july-correct-0081\",\"title\":\"AAI 2026: 6th Gen AMD EPYC Server CPUs Power the Agentic Data Center\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"GPU\",\"Intel\",\"PDF\",\"HTML\"],\"rank\":-11.191248974541931},{\"id\":\"historical-jan-apr-02\",\"title\":\"二、Google Cloud Next '26：AI Hypercomputer 与第八代 TPU 发布\",\"sourceType\":\"curated_item\",\"time\":\"2026-01_to_2026-04\",\"matchedTerms\":[\"GPU\",\"Intel\",\"KV\"],\"rank\":-9.688757721848434},{\"id\":\"runtime-569639751a0dbe7ef3ffccc5\",\"title\":\"[2608.17503] Predict Before Replay: Joint FEC and Flight Control for Reliable Scale-Up Links\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-18\",\"matchedTerms\":[\"PDF\",\"HTML\"],\"rank\":-8.032611962777274},{\"id\":\"runtime-bb3ff3b1181f90d6bd6ec57c\",\"title\":\"英伟达最强 Rubin GPU 架构发布，被 Blackwell 曝光\",\"sourceType\":\"ai_excellent_article\",\"time\":\"\",\"matchedTerms\":[\"GPU\",\"Intel\",\"HTML\"],\"rank\":-7.325140290534123}]"
AI摘要: "CacheBridge提出一种跨模型键值缓存迁移方法，通过架构索引映射、注意力对齐校准和有界映射器构建，避免接收模型重新预填充共享前缀。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-02T17:13:30.833Z"
采集批次: "2026年9月2日18点28分48秒"
采集批次ID: "20260902-182848-894"
去重键: "https://arxiv.org/abs/2609.00891"
---

## Computer Science > Artificial Intelligence

## Title:CacheBridge: Efficient Cross-Model KV Cache Transfer

Authors:[Xingyu Qu](https://arxiv.org/search/cs?searchtype=author&query=Qu,+X), [Siyuan Lu](https://arxiv.org/search/cs?searchtype=author&query=Lu,+S), [Zhiyu Chen](https://arxiv.org/search/cs?searchtype=author&query=Chen,+Z), [Sheng Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+S), [Tao Lin](https://arxiv.org/search/cs?searchtype=author&query=Lin,+T)

[View PDF](https://arxiv.org/pdf/2609.00891) [HTML (experimental)](https://arxiv.org/html/2609.00891v1)

> Abstract:Sharing context between LLMs in a multi-model system requires the receiving model to prefill the shared prefix because KV caches are model-specific. Recent closed-form cross-model KV transfer, hereafter Full-Head Mapping, avoids this replay by fitting a training-free affine mapper from source to target caches. However, its full-head design maps each target KV head from every source KV head in the selected layers, making transfer quality sensitive to architectural differences and causing mapper storage and application cost to grow with layer support. To this end, we introduce CacheBridge, which co-designs architecture-indexed mapper support, attention-aligned calibration, and bounded mapper construction while retaining a closed-form affine interface for online deployment. CacheBridge restricts each target head to a matched source head, weights reconstruction errors by causal attention sensitivity, and uses a fused GPU kernel to construct weighted sufficient statistics without materializing full observation tensors. Across three transfer directions, CacheBridge recovers the two Ministral 3 transfer directions where Full-Head Mapping loses substantial accuracy while preserving 99.83\\% mean target retention on Qwen3. On Qwen3 $14\mathrm{B}\to32\mathrm{B}$, it reduces mapper storage by $8\times$, accelerates application by up to $3.0\times$, matches \\fullhead with one tenth of the calibration data, and reduces 500-sequence construction from 92.63 to 8.63 seconds ($10.7\times$).

| Subjects: | Artificial Intelligence (cs.AI) |
| --- | --- |
| Cite as: | [arXiv:2609.00891](https://arxiv.org/abs/2609.00891) \[cs.AI\] |
|  | (or [arXiv:2609.00891v1](https://arxiv.org/abs/2609.00891v1) \[cs.AI\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.00891](https://doi.org/10.48550/arXiv.2609.00891) |

## Submission history

From: Siyuan Lu \[[view email](https://arxiv.org/show-email/007d7857/2609.00891)\]  
**\[v1\]** Tue, 1 Sep 2026 08:20:35 UTC (239 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.00891) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
