---
格式版本: 2
标题: "Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades"
原文链接: "https://arxiv.org/abs/2609.01345"
发布日期: "2026-09-01"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Tue, 1 Sep 2026 14:53:41 UTC (150 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-02T18:52:28+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-02T18:51:39+08:00"
入库时间: "2026-09-02T10:52:28.105Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=Scale-up&searchtype=all"
匹配关键词:
  - "Scale-up"
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 28
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是LLM推理级联、验证器盲区与纠错微调可靠性，并非超节点或机架级AI基础设施。当前页面是arXiv原创预印本摘要，新增了不同模型规模下盲区比例、升级率和真实错误率等实验结果；固定知识库未见同题重复，但未命中不能证明首次出现。文中没有机架拓扑、互连、供电、液冷、RAS或部署信息，且仅提供摘要。命中“应用与模型效率”硬否决项，不能进入业务情报库。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-09-02T18:53:40+08:00"
AI主题相关性: 0
AI来源权威性: 10
AI新颖性: 12
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 6
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.81"
AI评分知识库SHA256: "3b93d12e47749b3512f545f51c44c011bdc0931677c2cfe61e4df59a2b1a5a48"
AI评分知识库检索词: "[\"Scale-up\",\"Intel\",\"PDF\",\"arxiv.org/pdf/2609.01345\",\"HTML\",\"arxiv.org/html/2609.01345v1\",\"v1\",\"LLMs\",\"q_0\",\"beta_0\",\"arxiv.org/abs/2609.01345\",\"arxiv.org/abs/2609.01345v1\"]"
AI评分知识库命中: "[{\"id\":\"runtime-569639751a0dbe7ef3ffccc5\",\"title\":\"[2608.17503] Predict Before Replay: Joint FEC and Flight Control for Reliable Scale-Up Links\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-18\",\"matchedTerms\":[\"Scale-up\",\"PDF\",\"HTML\",\"v1\"],\"rank\":-13.013405860992304},{\"id\":\"july-correct-0111\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"PDF\",\"HTML\",\"v1\"],\"rank\":-11.543771882503226},{\"id\":\"july-correct-0081\",\"title\":\"AAI 2026: 6th Gen AMD EPYC Server CPUs Power the Agentic Data Center\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"Intel\",\"PDF\",\"HTML\"],\"rank\":-10.378395854671723},{\"id\":\"july-correct-0010\",\"title\":\"Schneider Electric and AMD release first Helios platform reference design to accelerate AI Factory deployment\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"PDF\",\"v1\"],\"rank\":-8.41606444929818},{\"id\":\"runtime-ffe0e6dc6f3bdd3f1ce31ae8\",\"title\":\"5000亿美元！英伟达押注AI基础设施 | SDNLAB | 专注网络创新技术\",\"sourceType\":\"ai_excellent_article\",\"time\":\"\",\"matchedTerms\":[\"Scale-up\",\"Intel\",\"HTML\"],\"rank\":-7.728228344047705}]"
AI摘要: "该研究测量了LLM推理级联中廉价验证器的可靠性成本，发现验证器盲点随学生模型能力增长而扩大（0.5B到32B时β从0.12升至0.55），在廉价学生加廉价验证器组合下最严重。用前沿验证器可将盲点降至约0.05，但导致近半查询升级；"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-02T17:13:27.032Z"
采集批次: "2026年9月2日18点28分48秒"
采集批次ID: "20260902-182848-894"
去重键: "https://arxiv.org/abs/2609.01345"
---

## Computer Science > Artificial Intelligence

## Title:Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades

Authors:[Dushyant Rajput](https://arxiv.org/search/cs?searchtype=author&query=Rajput,+D)

[View PDF](https://arxiv.org/pdf/2609.01345) [HTML (experimental)](https://arxiv.org/html/2609.01345v1)

> Abstract:Inference cascades cut cost by answering most queries with a cheap model and escalating a hard tail to a frontier model that acts as verifier. A natural extension closes the loop: fine-tune the cheap student on the verifier's rejections so the escalation rate, and cost, fall each round. We measure this loop on real LLMs and report four findings. First, the verifier's blind spot, the fraction of the student's wrong answers it accepts, is large and moves adversarially: it grows with student capability ($\beta$ from 0.12 to 0.55 as the student scales 0.5B to 32B) and shrinks with verifier capability, so it is worst in the cheap-student, cheap-verifier regime cascades exist to create. Second, buying it away returns the saving: a frontier verifier drives $\beta$ to about 0.05 but then escalates on 46% of hard-MATH queries against a 39% true error rate, paying the frontier price on nearly half of all traffic. Third, naive corrective fine-tuning on the verifier-rejected tail does not improve the small student but degrades and ultimately collapses it, across every teacher we tried (cross-family and same-family), so at this scale the self-improving loop is self-defeating. Fourth, through all of this the cascade's own dashboard, every metric computed through the verifier, reads a flat 3% error while true delivered error swings up to 32%: the system is blind to its own degradation by construction. We then give the theory that explains the blindness, a two-population conservation law, $\epsilon_\infty \lesssim q_0 \beta_0$, under which every in-loop metric improves while true quality does not, and a synthetic study that validates the mechanism. The practical conclusion: the reliability of a self-improving cascade cannot be read from any metric computed through its own verifier.

| Subjects: | Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Machine Learning (cs.LG) |
| --- | --- |
| Cite as: | [arXiv:2609.01345](https://arxiv.org/abs/2609.01345) \[cs.AI\] |
|  | (or [arXiv:2609.01345v1](https://arxiv.org/abs/2609.01345v1) \[cs.AI\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.01345](https://doi.org/10.48550/arXiv.2609.01345) |

## Submission history

From: Dushyant Rajput \[[view email](https://arxiv.org/show-email/61cff0a2/2609.01345)\]  
**\[v1\]** Tue, 1 Sep 2026 14:53:41 UTC (150 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.01345) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
