---
格式版本: 2
标题: "Avatar-Forever: Decoupled Parallel Training for High-Quality Real-Time Infinite Avatars"
原文链接: "https://arxiv.org/abs/2608.12107"
发布日期: "2026-08-12"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 12 Aug 2026 14:29:08 UTC (21,147 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-13T18:33:26+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-13T18:24:32+08:00"
入库时间: "2026-08-13T10:33:26.394Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=AI&searchtype=all"
匹配关键词:
  - "AI"
  - "GPU"
  - "throughput"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI质检状态: "评分失败"
AI评分尝试次数: 1
AI评分错误类型: "service_error"
AI评分错误: "LLM call failed; tried model chain: ali-deepseek-v4-flash -> tx-deepseek-v4-flash | Model ali-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131033313974729988268d9d6bmrh7ken)\",\"type\":\"new_api_error\"}} | Model tx-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131033338691487868268d9d6YSypLVg0)\",\"type\":\"new_api_error\"}}"
AI评分开始时间: "2026-08-13T10:33:29.876Z"
AI评分结束时间: "2026-08-13T10:33:34.028Z"
AI摘要: "研究者提出Avatar-Forever，一种面向高质量实时无限虚拟化身的解耦并行训练框架，将高效生成与长程鲁棒性分别并行训练，并引入ForeverCache减少流式推理中的冗余历史计算。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T03:40:18.497Z"
采集批次: "2026年8月13日18点24分27秒"
采集批次ID: "20260813-182427-1564e572"
去重键: "https://arxiv.org/abs/2608.12107"
---

## Computer Science > Computer Vision and Pattern Recognition

## Title:Avatar-Forever: Decoupled Parallel Training for High-Quality Real-Time Infinite Avatars

Authors:[Ruibin Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+R), [Tao Yang](https://arxiv.org/search/cs?searchtype=author&query=Yang,+T), [Zhiyuan Ma](https://arxiv.org/search/cs?searchtype=author&query=Ma,+Z), [Fangzhou Ai](https://arxiv.org/search/cs?searchtype=author&query=Ai,+F), [Shilei Wen](https://arxiv.org/search/cs?searchtype=author&query=Wen,+S), [Lei Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+L)

[View PDF](https://arxiv.org/pdf/2608.12107) [HTML (experimental)](https://arxiv.org/html/2608.12107v1)

> Abstract:Existing streaming video systems often rely on sequential, distillation-centered training pipelines to enable few-step long-video generation. However, this paradigm suffers from two limitations. First, failures or distribution shifts introduced in earlier stages affect later optimization, complicating the training process to converge. Second, the distillation-centric objective favours short-term generation but is prone to quality degradation when autoregressive errors accumulate over long rollouts. We propose Avatar-Forever, a decoupled parallel training framework for high-quality real-time infinite interactive avatars. Instead of coupling generation efficiency and long-horizon robustness under a sequential distillation pipeline, we treat them as two independent capabilities that can be trained in parallel. One branch performs full-parameter distillation to train an efficient generator with high visual quality, while another trains a lightweight long-horizon adapter via Recovery-oriented Rollout Training (RRT), which improves generation robustness under long-horizon inference conditions. Our decoupled parallel training design simplifies the overall training process and avoids unnecessary objective conflicts between few-step generation and long-horizon adaptation. We further introduce ForeverCache, a chunk-wise feature caching mechanism to substantially reduce redundant history computation during streaming inference. Built upon a 22B video foundation model, Avatar-Forever supports unbounded audio-driven avatar generation while maintaining identity consistency, motion coherence, and visual fidelity, enabling an end-to-end throughput of high-resolution 768x512 videos at 27.2 FPS on a single H100 GPU and providing a practical path toward stable digital humans.

| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| --- | --- |
| Cite as: | [arXiv:2608.12107](https://arxiv.org/abs/2608.12107) \[cs.CV\] |
|  | (or [arXiv:2608.12107v1](https://arxiv.org/abs/2608.12107v1) \[cs.CV\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.12107](https://doi.org/10.48550/arXiv.2608.12107) |

## Submission history

From: Ruibin Li \[[view email](https://arxiv.org/show-email/bb6b0a47/2608.12107)\]  
**\[v1\]** Wed, 12 Aug 2026 14:29:08 UTC (21,147 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.12107) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
