---
格式版本: 2
标题: "Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems"
原文链接: "https://arxiv.org/abs/2608.19140"
发布日期: "2026-08-19"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 19 Aug 2026 17:29:47 UTC (13 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-20T16:00:44+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-20T15:56:53+08:00"
入库时间: "2026-08-20T08:00:44.769Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=AI&searchtype=all"
匹配关键词:
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 21
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "该资料为arXiv人工智能论文，讨论模型输出精度度量方法，与超节点、AI Rack、机柜级AI基础设施、供电散热互连及量产落地等主题完全无关，属于明显无关内容。"
AI质检模型: "ali-deepseek-v4-flash"
AI质检时间: "2026-08-20T16:07:40+08:00"
AI主题相关性: 0
AI来源权威性: 8
AI新颖性: 5
AI技术细节: 3
AI商业部署信号: 0
AI完整性: 5
AI摘要: "作者提出前沿语言模型的比较应聚焦“精度”而非“能力”，即同一请求下输出围绕目标的集中程度，而非平均水平。该指标可通过固定任务多次运行低成本测得，并用于区分可修正的规则性失败与需更换模型的分散性失败；"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T03:20:06.288Z"
采集批次: "2026年8月20日14点19分32秒"
采集批次ID: "20260820-141932-079"
去重键: "https://arxiv.org/abs/2608.19140"
---

## Computer Science > Artificial Intelligence

## Title:Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems

Authors:[George Andrikopoulos](https://arxiv.org/search/cs?searchtype=author&query=Andrikopoulos,+G)

[View PDF](https://arxiv.org/pdf/2608.19140) [HTML (experimental)](https://arxiv.org/html/2608.19140v1)

> Abstract:Frontier language models are compared, marketed, and benchmarked on capability -- what their best or average output can achieve. I argue this measures the wrong axis. The models have saturated accuracy: their mean output lands on the target. What now separates one system from another in practice is precision: how tightly concentrated their outputs are around that target across repeated, identical requests. Borrowing the marksman's distinction, capability is where the average shot lands; reliability is the size of the group. I make three claims. First, precision, not capability, is the frontier differentiator between systems, and benchmark culture systematically fails to measure it, reporting central tendency rather than spread. Second, precision is measurable, cheaply and without circularity, by running a fixed suite of deterministically scored tasks many times at fixed temperature and computing the per-task consistency of outcomes -- no model-in-the-loop grader required. Third, the measurement is not merely descriptive but decision-guiding: it separates consistent failures (a tight group off-centre, correctable by the operating discipline of Paper 1 -- a sight adjustment) from scattered failures (a wide group, correctable only by changing the model or its sampling -- a rifle problem). I define a grouping metric, specify a harness, and show how tracking a human-AI pair's grouping over time yields the compounding signal that Paper 1's field study requires. A first real run, since replicated, illustrates both the method and its most important limit: one measured gap was closed completely by a single rule (0/5 -> 5/5), while a suite of tasks authored from the rules themselves found no value, because a frontier model already embodies explicit good practice -- establishing that a discipline's worth is found by measurement on real work, not constructed from its own rulebook.

| Comments: |  |
| --- | --- |
| Subjects: | Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Machine Learning (cs.LG); Software Engineering (cs.SE) |
| ACM classes: | D.2.8 |
| Cite as: | [arXiv:2608.19140](https://arxiv.org/abs/2608.19140) \[cs.AI\] |
|  | (or [arXiv:2608.19140v1](https://arxiv.org/abs/2608.19140v1) \[cs.AI\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.19140](https://doi.org/10.48550/arXiv.2608.19140) |

## Submission history

From: George Andrikopoulos \[[view email](https://arxiv.org/show-email/ed8b116c/2608.19140)\]  
**\[v1\]** Wed, 19 Aug 2026 17:29:47 UTC (13 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.19140) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
