---
格式版本: 2
标题: "Can Scene Text Recognition Read Rare Compositions?"
原文链接: "https://arxiv.org/abs/2609.00816"
发布日期: "2026-09-01"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Tue, 1 Sep 2026 07:16:42 UTC (2,492 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-02T18:53:44+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-02T18:51:39+08:00"
入库时间: "2026-09-02T10:53:44.373Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=Scale-up&searchtype=all"
匹配关键词:
  - "Scale-up"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 40
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是场景文字识别模型在罕见词与字符组合上的准确率退化，以及AR、CTC解码架构对比，完全不涉及超节点、AI机架或机架级基础设施。当前来源为作者提交的arXiv预印本原始摘要，尚非正式同行评审论文。固定知识库未发现同项研究，但Top 5不足以证明首次出现；正文新增了跨书写系统测试、扩容无效及CTC带来+2.5个百分点等实验结果，仅属于计算机视觉模型研究。没有客户、量产或部署信号，且页面只提供摘要而非全文。命中应用与模型效率类强否决，当前页面不值得作为超节点业务信息源。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-09-02T18:55:26+08:00"
AI主题相关性: 0
AI来源权威性: 10
AI新颖性: 14
AI技术细节: 9
AI商业部署信号: 0
AI完整性: 7
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.81"
AI评分知识库SHA256: "3b93d12e47749b3512f545f51c44c011bdc0931677c2cfe61e4df59a2b1a5a48"
AI评分知识库检索词: "[\"Scale-up\",\"NPU\",\"PDF\",\"arxiv.org/pdf/2609.00816\",\"HTML\",\"arxiv.org/html/2609.00816v1\",\"v1\",\"q3/q3\",\"CLIP4STR-Base\",\"CLIP4STR-Huge\",\"ViT-H/14\",\"LAION-2B\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0111\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"NPU\",\"PDF\",\"HTML\",\"v1\"],\"rank\":-17.581723725808548},{\"id\":\"runtime-569639751a0dbe7ef3ffccc5\",\"title\":\"[2608.17503] Predict Before Replay: Joint FEC and Flight Control for Reliable Scale-Up Links\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-18\",\"matchedTerms\":[\"Scale-up\",\"PDF\",\"HTML\",\"v1\"],\"rank\":-13.013405860992304},{\"id\":\"july-correct-0131\",\"title\":\"DeepSeek-V4如何在昇腾超节点高效完成全参数后训练？SLAI T-Rex技术报告解读\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"NPU\",\"PDF\",\"HTML\"],\"rank\":-10.795338416647796},{\"id\":\"july-correct-0088\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes - 智源社区论文\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"NPU\",\"HTML\"],\"rank\":-8.926562663983766},{\"id\":\"july-correct-0010\",\"title\":\"Schneider Electric and AMD release first Helios platform reference design to accelerate AI Factory deployment\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"PDF\",\"v1\"],\"rank\":-8.41606444929818}]"
AI摘要: "该研究重新评估场景文字识别，发现标准基准的高准确率掩盖了罕见词与罕见字符三元组组合上的显著性能下降，九个英文识别器在该组合上准确率比常规组合低10至18个百分点。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-02T17:13:16.413Z"
采集批次: "2026年9月2日18点28分48秒"
采集批次ID: "20260902-182848-894"
去重键: "https://arxiv.org/abs/2609.00816"
---

## Computer Science > Computer Vision and Pattern Recognition

## Title:Can Scene Text Recognition Read Rare Compositions?

Authors:[Genpei Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+G)

[View PDF](https://arxiv.org/pdf/2609.00816) [HTML (experimental)](https://arxiv.org/html/2609.00816v1)

> Abstract:Scene text recognition is reported as 89--97% accurate on the six standard benchmarks, and the problem is widely treated as saturated. We present an alternative reading. When the same test images are stratified jointly by ground-truth word rarity and character n-gram novelty against a reference corpus, accuracy at the rare-word x rare-trigram corner of the resulting 5x5 grid drops 10--18 pt below the q3/q3 centre across nine English specialised recognisers, and the same direction (corner below centre) holds on all 13 of 13 (language, model) pairs we test across four writing systems (Latin, Han, Han+kana, Arabic). The drop is not a capacity bottleneck. A 6x vision-backbone scale-up (CLIP4STR-Base 158M -> CLIP4STR-Huge 1.0B, OpenCLIP ViT-H/14 LAION-2B) leads every benchmark in aggregate accuracy yet leaves the stress corner unchanged (86.9 -> 86.5, within paired-bootstrap noise). Four converging probes--layer-wise probing, confidence-when-wrong, attention re-balancing, and a cross-script commit-vs-abstain error split--localise the failure to the autoregressive decoder's lexical prior. We then ask how much of the gap existing techniques recover. Of 16 non-architectural mitigations, the largest mean q5/q5 gain is +1.3 pt and none clears the paired-bootstrap noise floor; the only intervention that does is the architectural shift from autoregressive to CTC decoding (SVTRv2, +2.5 pt, p=0.02, n=474). A confidence-routed AR-CTC ensemble adds a directionally consistent +0.6 pt that stays within noise, and its dominant learned coefficient is each model's own minimum-softmax confidence--independently echoing the mechanism above. No configuration we test improves both the compositional corner and aggregate accuracy. The rare-input long tail thus points to architectural change rather than added capacity.

| Comments: |  |
| --- | --- |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | [arXiv:2609.00816](https://arxiv.org/abs/2609.00816) \[cs.CV\] |
|  | (or [arXiv:2609.00816v1](https://arxiv.org/abs/2609.00816v1) \[cs.CV\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.00816](https://doi.org/10.48550/arXiv.2609.00816) |

## Submission history

From: Genpei Zhang \[[view email](https://arxiv.org/show-email/002762e3/2609.00816)\]  
**\[v1\]** Tue, 1 Sep 2026 07:16:42 UTC (2,492 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.00816) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
