---
格式版本: 2
标题: "FleetSieve: Decision-Critical Profiling for SLO-Aware LLM Fleet Configuration"
原文链接: "https://arxiv.org/abs/2608.19659"
发布日期: "2026-08-20"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Thu, 20 Aug 2026 05:54:29 UTC (13 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-21T16:19:12+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-21T16:17:02+08:00"
入库时间: "2026-08-21T08:19:12.823Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=Oracle&searchtype=all"
匹配关键词:
  - "GPU"
  - "performance"
  - "latency"
相关厂家:
  - "Oracle"
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 25
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "论文主题为LLM服务集群配置优化，属推理调度与性能剖析，不涉及超节点、AI Rack、机柜级硬件、供电散热或高速互连等核心关注领域，与项目主题明显无关。"
AI质检模型: "ali-deepseek-v4-flash"
AI质检时间: "2026-08-21T16:19:21+08:00"
AI主题相关性: 0
AI来源权威性: 10
AI新颖性: 5
AI技术细节: 5
AI商业部署信号: 0
AI完整性: 5
AI摘要: "FleetSieve提出一种面向LLM服务集群配置的决策关键型剖析方法，仅测量对资源分配决策有影响的配置，以降低GPU开销。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T02:12:23.544Z"
采集批次: "2026年8月21日13点55分53秒"
采集批次ID: "20260821-135553-743"
去重键: "https://arxiv.org/abs/2608.19659"
---

## Computer Science > Machine Learning

## Title:FleetSieve: Decision-Critical Profiling for SLO-Aware LLM Fleet Configuration

Authors:[Huang Cheng](https://arxiv.org/search/cs?searchtype=author&query=Cheng,+H), [Scott Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+S), [Aubert Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+A)

[View PDF](https://arxiv.org/pdf/2608.19659) [HTML (experimental)](https://arxiv.org/html/2608.19659v1)

> Abstract:Choosing tensor-parallel (TP) degrees and replica counts for an LLM serving fleet is difficult because performance is not monotonic in TP and the feasible choice can change with load. Exhaustive profiling resolves this uncertainty, but measures many configurations that do not affect the final resource allocation. We present FleetSieve, which selects measurements according to their expected effect on a resource-coupled, SLO-aware fleet decision. FleetSieve models capacity and tail latency jointly, compares conservative and optimistic allocations, and stops when their remaining decision gap is below a specified tolerance. On a fixed H100 measurement grid for a 31B-parameter open-weight model, FleetSieve reaches the oracle aggregate decision using 22,200 GPU-seconds, 6.9% less than uniform random profiling in the fixed comparison. Across 200 random reveal orders, its mean saving over random profiling is 5.4% (95% bootstrap CI: 3.5-7.2%). The fixed-comparison saving is 21.5% for Chat, while FleetSieve does not use the fewest GPU-seconds for Code. Joint capacity and tail modeling also avoids selecting a configuration whose 46.4-second completion p99 violates a 30-second SLO. In a 16-GPU allocation, an incorrect sparse-profile decision loses up to 1.93 requests/s and 12.4 percentage points of max-min fulfillment. Boundary repeats and BurstGPT measurements support the observed load-dependent tail-latency mechanism.

| Subjects: | Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC) |
| --- | --- |
| Cite as: | [arXiv:2608.19659](https://arxiv.org/abs/2608.19659) \[cs.LG\] |
|  | (or [arXiv:2608.19659v1](https://arxiv.org/abs/2608.19659v1) \[cs.LG\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.19659](https://doi.org/10.48550/arXiv.2608.19659) |

## Submission history

From: Huang Cheng \[[view email](https://arxiv.org/show-email/3c688ceb/2608.19659)\]  
**\[v1\]** Thu, 20 Aug 2026 05:54:29 UTC (13 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.19659) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
