---
格式版本: 2
标题: "Zero-Shot SAM2 Segmentation and Vision Transformer-Based Recognition of Elamite Cuneiform Symbols from Degraded Tablet Images"
原文链接: "https://arxiv.org/abs/2608.18544"
发布日期: "2026-08-19"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 19 Aug 2026 05:02:41 UTC (22,768 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-20T14:34:03+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-20T14:32:53+08:00"
入库时间: "2026-08-20T06:34:03.506Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=NVIDIA&searchtype=all"
匹配关键词:
  - "GPU"
  - "performance"
相关厂家:
  - "NVIDIA"
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 8
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "该论文为计算机视觉领域的楔形文字识别研究，仅提及NVIDIA A100作为运行设备，与超节点/AI Rack/机柜级AI基础设施完全无关。"
AI质检模型: "ali-deepseek-v4-flash"
AI质检时间: "2026-08-20T14:34:22+08:00"
AI主题相关性: 0
AI来源权威性: 5
AI新颖性: 2
AI技术细节: 1
AI商业部署信号: 0
AI完整性: 0
AI摘要: "研究人员提出EpigraphNet，结合零样本SAM2分割与Vision Transformer，对Persepolis Fortification Archive的1,239张退化泥板图像进行Elamite楔形文字识别。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T03:19:57.172Z"
采集批次: "2026年8月20日14点19分32秒"
采集批次ID: "20260820-141932-079"
去重键: "https://arxiv.org/abs/2608.18544"
---

## Computer Science > Computer Vision and Pattern Recognition

## Title:Zero-Shot SAM2 Segmentation and Vision Transformer-Based Recognition of Elamite Cuneiform Symbols from Degraded Tablet Images

[View PDF](https://arxiv.org/pdf/2608.18544) [HTML (experimental)](https://arxiv.org/html/2608.18544v1)

> Abstract:Automated recognition of ancient cuneiform script poses a compound signal-degradation problem: the three-dimensional relief of clay tablets creates spatially varying illumination and cast shadows, surface erosion introduces structured noise that overlaps with genuine sign impressions, and severe class imbalance across 141 sign categories undermines classifier reliability. We introduce EpigraphNet, a segmentation-guided transformer pipeline evaluated on the Persepolis Fortification Archive. From 1,239 annotated tablet images, brightness-adaptive morphological preprocessing and zero-shot SAM2-Large segmentation generate clean binary symbol masks, which a fine-tuned Vision Transformer (ViT-B/16) with inverse-frequency class weighting then classifies. EpigraphNet reaches 86.41% top-1 accuracy on a 132-class benchmark, a 17.21 percentage-point gain over the strongest CNN baseline (ResNet-101, 69.20%) and 5.31-12.91% over four modern backbones (DeiT-B/16, Swin-B, ConvNeXt-B, EfficientNet-B4) under identical conditions. The full pipeline runs at approximately 18 ms per sign on an NVIDIA A100 GPU. A lower Spearman correlation between sign frequency and per-class performance indicates more balanced recognition across frequent and rare classes. Implementation is available at: [this http URL](http://github.com/r11up/sam-guided-vit)

| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC) |
| --- | --- |
| Cite as: | [arXiv:2608.18544](https://arxiv.org/abs/2608.18544) \[cs.CV\] |
|  | (or [arXiv:2608.18544v1](https://arxiv.org/abs/2608.18544v1) \[cs.CV\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.18544](https://doi.org/10.48550/arXiv.2608.18544) |

## Submission history

From: Utsav Poudel \[[view email](https://arxiv.org/show-email/31b5ac79/2608.18544)\]  
**\[v1\]** Wed, 19 Aug 2026 05:02:41 UTC (22,768 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.18544) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
