---
格式版本: 2
标题: "mzCache: On-Device LLM Memory Management under Multitasking"
原文链接: "https://arxiv.org/abs/2609.01338"
发布日期: "2026-09-01"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Tue, 1 Sep 2026 14:49:20 UTC (1,007 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-02T18:52:05+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-02T18:51:36+08:00"
入库时间: "2026-09-02T10:52:05.350Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=GPU&searchtype=all"
匹配关键词:
  - "GPU"
  - "latency"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 35
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是移动设备多任务环境下的端侧LLM内存管理，提出mzCache并在Android应用中实现，通过细粒度共享缓冲区、混合交换和反向淘汰将TTFT降低2.1—5.5倍；并非超节点、AI Rack或机架级基础设施。来源为arXiv论文摘要页，学术来源可追溯但当前正文缺少论文全文细节。固定知识库未发现mzCache同题记录，但Top 5不完整，不能据此认定首次出现；可确认的新增仅是本文所述端侧原型及实验结果。无客户、量产或生产部署信号，且机制面向移动SoC和单设备推理，不能形成可复用的机架级通信、RAS、调度或内存基础设施能力，命中应用与模型效率及非生产学术原型否决项。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-09-02T18:52:39+08:00"
AI主题相关性: 1
AI来源权威性: 11
AI新颖性: 12
AI技术细节: 4
AI商业部署信号: 0
AI完整性: 7
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.81"
AI评分知识库SHA256: "3b93d12e47749b3512f545f51c44c011bdc0931677c2cfe61e4df59a2b1a5a48"
AI评分知识库检索词: "[\"GPU\",\"LLM\",\"PDF\",\"arxiv.org/pdf/2609.01338\",\"HTML\",\"arxiv.org/html/2609.01338v1\",\"KV\",\"CPU-side\",\"URL\",\"arxiv.org/abs/2609.01338\",\"arxiv.org/abs/2609.01338v1\",\"doi.org/10.48550/arXiv.2609.01338\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0111\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"LLM\",\"PDF\",\"HTML\",\"URL\"],\"rank\":-10.814912173575701},{\"id\":\"runtime-569639751a0dbe7ef3ffccc5\",\"title\":\"[2608.17503] Predict Before Replay: Joint FEC and Flight Control for Reliable Scale-Up Links\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-18\",\"matchedTerms\":[\"PDF\",\"HTML\",\"URL\"],\"rank\":-9.85203093015941},{\"id\":\"runtime-35f284698f34b54313c05e16\",\"title\":\"7 Restore LLM Inference Capacity in Seconds with Shadow Engine Recovery in NVIDIA Dynamo | August 2026 NVIDIA Dynamo's shadow engine recovery feature, presented by Maksim Khadkevich, Vikram Sharma Mailthody, Mohammed Abdulwahhab, and Schwinn Saereesitthipitak, reduces LLM inference recovery time from 283 seconds to 7.3 seconds by utilizing a preinitialized shadow engine on the same GPUs as the active engine. It leverages NVIDIA CUDA, GPU Memory Service (GMS), and Dynamic Resource Allocation (DRA\",\"sourceType\":\"ai_excellent_article\",\"time\":\"\",\"matchedTerms\":[\"GPU\",\"LLM\",\"HTML\",\"KV\"],\"rank\":-8.884482799170714},{\"id\":\"july-correct-0117\",\"title\":\"ODCC分享 | 华为严可荣：OCS全光超节点集群技术研究分析\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"LLM\",\"HTML\",\"KV\"],\"rank\":-8.58049846646998},{\"id\":\"july-correct-0090\",\"title\":\"Scaling Agentic AI Factories Through Extreme Co-Design with NVIDIA BlueField | NVIDIA Technical Blog\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"GPU\",\"LLM\",\"KV\",\"URL\"],\"rank\":-7.906458738150938}]"
AI摘要: "mzCache 提出面向移动端多任务场景的 LLM 推理内存管理系统，通过弹性驱逐和移动 SoC 统一内存实现 GPU 零等待推理及 CPU 侧并发恢复。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-02T17:13:31.680Z"
采集批次: "2026年9月2日18点28分48秒"
采集批次ID: "20260902-182848-894"
去重键: "https://arxiv.org/abs/2609.01338"
---

## Computer Science > Operating Systems

## Title:mzCache: On-Device LLM Memory Management under Multitasking

Authors:[Hongseung Yu](https://arxiv.org/search/cs?searchtype=author&query=Yu,+H), [Minsung Kim](https://arxiv.org/search/cs?searchtype=author&query=Kim,+M), [Jongseok Park](https://arxiv.org/search/cs?searchtype=author&query=Park,+J), [Kyunghan Lee](https://arxiv.org/search/cs?searchtype=author&query=Lee,+K)

[View PDF](https://arxiv.org/pdf/2609.01338) [HTML (experimental)](https://arxiv.org/html/2609.01338v1)

> Abstract:On-device mobile Large Language Model (LLM) inference is gaining significant attention. However, mobile devices operate in highly dynamic multitasking environments where users frequently switch between applications. This creates memory pressure, forcing LLM memory (model weights and KV cache) to be evicted by the operating system. When a new inference request arrives, the inference system must restore the evicted memory through slow storage reads or recompute the entire KV cache, severely degrading responsiveness. To address this, we present mzCache, an on-device LLM inference system with specialized memory management for multitasking environments. Under unpredictable memory pressure, mzCache elastically evicts LLM memory and leverages the unified memory of mobile SoCs to enable zero-wait inference on the GPU with concurrent CPU-side restoration. mzCache realizes this through restoration-oriented memory management: LLM memory is partitioned into fine-grained shared buffers to enable partial eviction and restoration with concurrent cross-processor access, while hybrid swap and backward-out eviction policies ensure low-latency restoration from any eviction state. Implemented on [this http URL](http://llama.cpp/) and deployed as an Android application, mzCache achieves 2.1-5.5 $\times$ reduction in Time-to-First-Token compared to storage-backed partial offload and demonstrates its effectiveness in real multitasking scenarios.

| Comments: |  |
| --- | --- |
| Subjects: | Operating Systems (cs.OS); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG) |
| Cite as: | [arXiv:2609.01338](https://arxiv.org/abs/2609.01338) \[cs.OS\] |
|  | (or [arXiv:2609.01338v1](https://arxiv.org/abs/2609.01338v1) \[cs.OS\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.01338](https://doi.org/10.48550/arXiv.2609.01338) |
| Related DOI: | [https://doi.org/10.1145/3795866.3844495](https://doi.org/10.1145/3795866.3844495) |

## Submission history

From: Hongseung Yu \[[view email](https://arxiv.org/show-email/2bbc2f44/2609.01338)\]  
**\[v1\]** Tue, 1 Sep 2026 14:49:20 UTC (1,007 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.01338) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
