---
格式版本: 2
标题: "Pair-Level Essay-Scale Republication and Reuse from Fragmented Historical Text Reuse: A Workflow Study on Eighteenth-Century Books and Newspapers"
原文链接: "https://arxiv.org/abs/2608.27343"
发布日期: "2026-08-27"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Thu, 27 Aug 2026 16:43:27 UTC (101 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-30T15:12:31+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-30T15:11:38+08:00"
入库时间: "2026-08-30T07:12:31.886Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=Scale-up&searchtype=all"
匹配关键词:
  - "Scale-up"
  - "deployment"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 19
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是18世纪书籍与报纸文本复用的证据整合研究，与超节点、AI机架及关键基础设施无关。来源为可追溯的arXiv论文摘要页，但仅提供摘要而非全文。知识库未见同项研究，不能据此认定首次出现；正文新增的规则流程、F1及文本复用审计结果也不属于业务所需事实，无机架级技术细节或商业部署信号，命中弱相关强否决。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-08-30T15:12:42+08:00"
AI主题相关性: 0
AI来源权威性: 11
AI新颖性: 2
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 6
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.55"
AI评分知识库SHA256: "ca56afec1ff616b1fa4394cd1d08b1e2a6af47eccbe70c802c9f1fea6037e803"
AI评分知识库检索词: "[\"Scale-up\",\"NPU\",\"C3\",\"A4kel\",\"A4\",\"PDF\",\"arxiv.org/pdf/2608.27343\",\"HTML\",\"arxiv.org/html/2608.27343v1\",\"ECCO\",\"LLM\",\"F1\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0111\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"NPU\",\"PDF\",\"HTML\",\"LLM\"],\"rank\":-14.493749653978089},{\"id\":\"july-correct-0131\",\"title\":\"DeepSeek-V4如何在昇腾超节点高效完成全参数后训练？SLAI T-Rex技术报告解读\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"NPU\",\"C3\",\"A4\",\"PDF\",\"HTML\",\"LLM\",\"F1\"],\"rank\":-11.988553635560805},{\"id\":\"runtime-569639751a0dbe7ef3ffccc5\",\"title\":\"[2608.17503] Predict Before Replay: Joint FEC and Flight Control for Reliable Scale-Up Links\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-18\",\"matchedTerms\":[\"Scale-up\",\"PDF\",\"HTML\",\"F1\"],\"rank\":-10.481970259177816},{\"id\":\"july-correct-0088\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes - 智源社区论文\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"NPU\",\"HTML\",\"LLM\"],\"rank\":-10.465171912639056},{\"id\":\"july-correct-0001\",\"title\":\"全球首颗2nm GPU来了！苏姿丰甩出“最强AI机架”，CPU性能干翻英伟达 - 智东西\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"Scale-up\",\"NPU\",\"HTML\",\"LLM\"],\"rank\":-7.132080386931695}]"
AI摘要: "该研究提出了一种对碎片化文本重用证据进行配对级整合的工作流，用于恢复18世纪书籍与报纸中休谟文章的再版和重用。在标注切片上该方法F1达0.948，人工审计确认全部176个报纸预测均为真实重用案例。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-08-30T18:08:57.827Z"
采集批次: "2026年8月30日14点11分37秒"
采集批次ID: "20260830-141137-130"
去重键: "https://arxiv.org/abs/2608.27343"
---

## Computer Science > Computation and Language

## Title:Pair-Level Essay-Scale Republication and Reuse from Fragmented Historical Text Reuse: A Workflow Study on Eighteenth-Century Books and Newspapers

Authors:[Ke Shu](https://arxiv.org/search/cs?searchtype=author&query=Shu,+K), [Kira Hinderks](https://arxiv.org/search/cs?searchtype=author&query=Hinderks,+K), [Eetu Mäkelä](https://arxiv.org/search/cs?searchtype=author&query=M%C3%A4kel%C3%A4,+E), [Mikko Tolonen](https://arxiv.org/search/cs?searchtype=author&query=Tolonen,+M)

[View PDF](https://arxiv.org/pdf/2608.27343) [HTML (experimental)](https://arxiv.org/html/2608.27343v1)

> Abstract:This paper addresses the recovery of essay-scale republication and reuse from fragmented text-reuse evidence, a setting whose central challenge is pair-level evidence consolidation and not fragment retrieval alone. The study focuses on a candidate set centered on essays by eighteenth-century Scottish philosopher David Hume, spanning books from ECCO (Eighteenth Century Collections Online) and historical newspapers. Because the input consists of fragmented reuse hits instead of clean document pairs, and positive coverage is inherently incomplete, we formulate the task as pair-level evidence consolidation into plausible transmission relations and compare three methodological families: a staged rule-based workflow, baselines (a decision tree and two direct LLM settings), and automated rule adaptation. On labeled ECCO--ECCO slices, pair-level feature aggregation alone already reaches 0.948 F1 on the main labeled slice, while the final workflow gives the strongest overall precision-recall trade-off among the tested rule stages. On the full ECCO--ECCO candidate universe, direct LLM baselines flag up to 14,886 pairs as reprints compared to 771 for the final workflow, behaving in this direct-prompt setup as high-recall candidate expanders rather than precision-controlled deployment classifiers. On ECCO--Newspaper, manual audit confirms all 176 predicted positives as genuine cases of republication or reuse, while issue duplication and source-side multiplicity reveal additional provenance structure. Under incomplete ground truth, auditable pair-level evidence consolidation provides a practical way to produce compact candidate spaces for historical inspection.

| Subjects: | Computation and Language (cs.CL) |
| --- | --- |
| Cite as: | [arXiv:2608.27343](https://arxiv.org/abs/2608.27343) \[cs.CL\] |
|  | (or [arXiv:2608.27343v1](https://arxiv.org/abs/2608.27343v1) \[cs.CL\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.27343](https://doi.org/10.48550/arXiv.2608.27343) |
| Related DOI: | [https://doi.org/10.1145/3799682.3839901](https://doi.org/10.1145/3799682.3839901) |

## Submission history

From: Ke Shu \[[view email](https://arxiv.org/show-email/a34de6a9/2608.27343)\]  
**\[v1\]** Thu, 27 Aug 2026 16:43:27 UTC (101 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.27343) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
