---
格式版本: 2
标题: "Competence, Not Accuracy: A Diagnostic for Reference-Free Judge Gates in Skill Optimization"
原文链接: "https://arxiv.org/abs/2608.18719"
发布日期: "2026-08-19"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 19 Aug 2026 09:19:23 UTC (1,820 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-20T15:22:38+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-20T15:18:01+08:00"
入库时间: "2026-08-20T07:22:38.912Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=deployment&searchtype=all"
匹配关键词:
  - "deployment"
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 7
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "该资料为arXiv学术论文，主题是LLM judge gate诊断，与超节点、AI Rack、机柜级AI基础设施完全无关，仅因搜索词deployment命中，无有效技术或商业信息。"
AI质检模型: "ali-deepseek-v4-flash"
AI质检时间: "2026-08-20T15:26:31+08:00"
AI主题相关性: 0
AI来源权威性: 5
AI新颖性: 0
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 2
AI摘要: "该研究提出一种低成本的部署前诊断方法，用于判断技能优化中的无参考LLM评判门是否具备区分正确与错误答案的判别力。作者将评判门形式化为潜在求解器，给出鉴别力下界及必要条件，并实验发现基准准确率会高估其实际能力。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T03:19:16.268Z"
采集批次: "2026年8月20日14点19分32秒"
采集批次ID: "20260820-141932-079"
去重键: "https://arxiv.org/abs/2608.18719"
---

## Computer Science > Artificial Intelligence

## Title:Competence, Not Accuracy: A Diagnostic for Reference-Free Judge Gates in Skill Optimization

Authors:[Chenle Chen](https://arxiv.org/search/cs?searchtype=author&query=Chen,+C), [Yangbo Wei](https://arxiv.org/search/cs?searchtype=author&query=Wei,+Y), [Chao Yao](https://arxiv.org/search/cs?searchtype=author&query=Yao,+C), [Shaoqiang Lu](https://arxiv.org/search/cs?searchtype=author&query=Lu,+S), [Junhong Qian](https://arxiv.org/search/cs?searchtype=author&query=Qian,+J), [Chen Wu](https://arxiv.org/search/cs?searchtype=author&query=Wu,+C), [Lei He](https://arxiv.org/search/cs?searchtype=author&query=He,+L)

[View PDF](https://arxiv.org/pdf/2608.18719) [HTML (experimental)](https://arxiv.org/html/2608.18719v1)

> Abstract:Text-space skill optimization adapts a frozen agent by evolving a natural-language skill document, accepting each candidate through a validation gate. Existing gates rely on verifiable rewards, confining these methods to tasks with an automatic verifier. Replacing the verifier with an LLM-judge gate would lift that restriction, but whether such a gate carries usable signal is untested. We ask a prior question: can we tell, before placing a judge in the loop, whether its scores separate correct from incorrect answers at all? We formalize a reference-free judge as a latent solver -- its verdict rests on agreement with whatever it would itself conclude, so its capacity to evaluate is bounded by its capacity to solve. The model yields a closed-form bound on discriminability (ROC-AUC) in the judge's competence $c$ and answer-space size $k$, a necessary condition $c > 1/k$, and the result that the marginal AUC is confounded by item difficulty while a within-question estimator is not. A non-intervening probe records judge scores on genuine optimization runs without altering any decision. We find discriminability at chance where competence sits near the floor and usable above it; that a judge's benchmark accuracy overstates the competence that matters; and, in a closed-loop study, that the screen predicts which kind of gating error occurs. The result is a cheap pre-deployment diagnostic for judge gates.

| Comments: |  |
| --- | --- |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | [arXiv:2608.18719](https://arxiv.org/abs/2608.18719) \[cs.AI\] |
|  | (or [arXiv:2608.18719v1](https://arxiv.org/abs/2608.18719v1) \[cs.AI\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.18719](https://doi.org/10.48550/arXiv.2608.18719) |

## Submission history

From: Chenle Chen \[[view email](https://arxiv.org/show-email/825f9c50/2608.18719)\]  
**\[v1\]** Wed, 19 Aug 2026 09:19:23 UTC (1,820 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.18719) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
