---
格式版本: 2
标题: "User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling"
原文链接: "https://arxiv.org/abs/2608.11840"
发布日期: "2026-08-12"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 12 Aug 2026 09:28:43 UTC (2,610 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-13T18:34:10+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-13T18:24:32+08:00"
入库时间: "2026-08-13T10:34:11.115Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=AI&searchtype=all"
匹配关键词:
  - "AI"
  - "performance"
  - "latency"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI质检状态: "评分失败"
AI评分尝试次数: 1
AI评分错误类型: "service_error"
AI评分错误: "LLM call failed; tried model chain: ali-deepseek-v4-flash -> tx-deepseek-v4-flash | Model ali-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131034164156918888268d9d64n5AU73c)\",\"type\":\"new_api_error\"}} | Model tx-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131034175647243338268d9d6u5b9Aem5)\",\"type\":\"new_api_error\"}}"
AI评分开始时间: "2026-08-13T10:34:12.146Z"
AI评分结束时间: "2026-08-13T10:34:17.690Z"
AI摘要: "该文提出一种用户协助的协作分布式推理系统，将专用基础设施与用户自愿贡献资源结合，以在不等比扩展集中设施的情况下满足AI推理需求。模拟显示，随用户规模增长，分布式调度可显著改善请求完成率和P99延迟，并降低专用资源消耗。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T03:40:24.642Z"
采集批次: "2026年8月13日18点24分27秒"
采集批次ID: "20260813-182427-1564e572"
去重键: "https://arxiv.org/abs/2608.11840"
---

## Computer Science > Distributed, Parallel, and Cluster Computing

## Title:User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling

[View PDF](https://arxiv.org/pdf/2608.11840) [HTML (experimental)](https://arxiv.org/html/2608.11840v1)

> Abstract:Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise with demand. We propose a collaborative distributed inference system combining dedicated infrastructure with resources contributed by service users. Dedicated resources provide baseline capacity for maintaining quality of service (QoS), while volunteered resources absorb increasing demand without proportional growth in centralized infrastructure. To capture stochastic and dynamic interactions among users, resources, tasks, and policies, we develop a high-dimensional generative Markov model with structured temporal factorization. The model supports simulation and provides a foundation for task scheduling and QoS-aware resource allocation optimization. We evaluate the system across user populations, resource capacities, and centralized and distributed scheduling policies. Simulations show that distributed scheduling becomes increasingly advantageous as the user population grows, improving request completion and P99 latency while substantially reducing dedicated resource consumption. These results demonstrate the feasibility of user-assisted collaborative inference for infrastructure-efficient autoscaling.

| Subjects: | Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Networking and Internet Architecture (cs.NI); Performance (cs.PF) |
| --- | --- |
| Cite as: | [arXiv:2608.11840](https://arxiv.org/abs/2608.11840) \[cs.DC\] |
|  | (or [arXiv:2608.11840v1](https://arxiv.org/abs/2608.11840v1) \[cs.DC\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.11840](https://doi.org/10.48550/arXiv.2608.11840) |

## Submission history

From: Alfreds Lapkovskis \[[view email](https://arxiv.org/show-email/1b879912/2608.11840)\]  
**\[v1\]** Wed, 12 Aug 2026 09:28:43 UTC (2,610 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.11840) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
