---
格式版本: 2
标题: "PersianAnonymizer: Evaluating LLM-Labeled Training for Efficient NER-based Anonymization in Persian"
原文链接: "https://arxiv.org/abs/2609.00958"
发布日期: "2026-09-01"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "llm:scrape:strict_html_body"
发布时间证据: "div class=dateline: [Submitted on 1 Sep 2026]"
发布时间校准原因: "arXiv提交日期明确标注为1 Sep 2026，属于文章发布时间；LREC会议日期为会议举办时间，应排除。"
发布时间校准置信度: "1"
发布时间候选数量: 24
发布时间严格候选数量: 8
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-02T18:51:58+08:00"
发布时间仲裁状态: "confirmed"
发布时间仲裁尝试次数: 1
发布时间仲裁耗时毫秒: 17878
发现时间: "2026-09-02T18:51:36+08:00"
入库时间: "2026-09-02T10:52:16.305Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=GPU&searchtype=all"
匹配关键词:
  - "GPU"
  - "deployment"
  - "latency"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 25
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是利用LLM标注训练波斯语NER模型以匿名化客户聊天，H200与RTX 3090仅用于比较标注延迟，不涉及机架级AI架构、互连、供电、散热或RAS。来源为arXiv论文摘要页并附LREC 2026出版信息，学术来源可追溯但正文并非完整论文。固定知识库未发现该NER实验的历史记录，但新增的语料、F1/LCR和约2分钟处理4万条消息等结果属于单一语言应用与模型效率，不构成超节点业务新增；无机架产品、客户采购、量产或基础设施部署信号。命中“应用与模型效率”硬否决项。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-09-02T18:53:15+08:00"
AI主题相关性: 0
AI来源权威性: 10
AI新颖性: 6
AI技术细节: 2
AI商业部署信号: 1
AI完整性: 6
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.81"
AI评分知识库SHA256: "3b93d12e47749b3512f545f51c44c011bdc0931677c2cfe61e4df59a2b1a5a48"
AI评分知识库检索词: "[\"GPU\",\"RAS\",\"LLM-Labeled\",\"NER-based\",\"PDF\",\"arxiv.org/pdf/2609.00958\",\"NER\",\"LLMs\",\"DeepSeek-V3-0324\",\"GPT-OSS-120B\",\"OSS\",\"Qwen3-235B-A22B-Instruct-2507\"]"
AI评分知识库命中: "[{\"id\":\"runtime-dccd09f98e2ca028070793f0\",\"title\":\"OCI Achieves NVIDIA Exemplar Cloud Validation for NVIDIA GB300 NVL72 and HGX B300 | cloud-infrastructure\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-22\",\"matchedTerms\":[\"GPU\",\"RAS\",\"NER\",\"OSS\"],\"rank\":-11.517728513634218},{\"id\":\"runtime-5581f72bc66fda8e6c6bf711\",\"title\":\"Advancing Standards-Based AI Fabric RAS Through COSMOS Integration\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-10\",\"matchedTerms\":[\"GPU\",\"RAS\",\"PDF\",\"NER\",\"OSS\"],\"rank\":-7.329850491814313},{\"id\":\"runtime-bf6a5bdbe28961c50fb91e45\",\"title\":\"Amazon EC2 P6-B300 instances are now available in additional AWS Regions\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-28\",\"matchedTerms\":[\"GPU\",\"LLMs\"],\"rank\":-7.299574186623396},{\"id\":\"july-correct-0112\",\"title\":\"超前点映AMD Advancing AI 2026：AMD AI YES？-36氪\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"GPU\",\"OSS\"],\"rank\":-6.3070723689970585},{\"id\":\"historical-jun-010\",\"title\":\"爱建证券-电子行业专题报告：Vera Rubin量产提速，RTX Spark打开终端AI新空间-260608.pdf\",\"sourceType\":\"curated_item\",\"time\":\"2026-06\",\"matchedTerms\":[\"PDF\"],\"rank\":-5.954285493762395}]"
AI摘要: "PersianAnonymizer 利用三个指令微调 LLM 生成标注数据，训练紧凑的 MatinaRoberta 命名实体识别模型，用于波斯语客户聊天记录匿名化。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-02T17:13:28.198Z"
采集批次: "2026年9月2日18点28分48秒"
采集批次ID: "20260902-182848-894"
去重键: "https://arxiv.org/abs/2609.00958"
---

## Computer Science > Computation and Language

## Title:PersianAnonymizer: Evaluating LLM-Labeled Training for Efficient NER-based Anonymization in Persian

[View PDF](https://arxiv.org/pdf/2609.00958)

> Abstract:We target practical anonymization of Persian customer chats by training a compact NER model from LLM-labeled supervision and selecting the best labeler for deployment. We compare three instruction-tuned LLMs: DeepSeek-V3-0324, GPT-OSS-120B, and Qwen3-235B-A22B-Instruct-2507, to produce span annotations under a shared JSON protocol, yielding four corpora (OSS\_ZeroShot, Qwen\_ZeroShot, Qwen\_FewShot, DeepSeek\_FewShot). A MatinaRoberta-based token-classifier is trained per corpus and evaluated with token-level Precision/Recall/F1 (overall and per-class). We also report Label Coverage Recall (LCR), the proportion of gold non-O tokens predicted as non-O, and quantify cross-labeler behavior via a token-level Venn on test annotations. Finally, we contrast test-set annotation latency of the LLMs on H200 nodes with the trained NER's test-time labeling on a single RTX 3090. Results show that supervision from OSS\_ZeroShot yields the strongest macro-F1 and LCR, while the resulting NER labels an entire 40K-message test set in approximately 2 minutes on one consumer GPU. This establishes a practical path to high-quality, low-cost anonymization for Persian industrial data.

| Comments: |  |
| --- | --- |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | [arXiv:2609.00958](https://arxiv.org/abs/2609.00958) \[cs.CL\] |
|  | (or [arXiv:2609.00958v1](https://arxiv.org/abs/2609.00958v1) \[cs.CL\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.00958](https://doi.org/10.48550/arXiv.2609.00958) |
| Journal reference: | Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026), pages 4497-4506, 11-16 May 2026. ELRA Language Resources Association (ELRA), 2026 |
| Related DOI: | [https://doi.org/10.63317/57u2ica9225o](https://doi.org/10.63317/57u2ica9225o) |

## Submission history

From: Mohammad Hossein Shalchian \[[view email](https://arxiv.org/show-email/f1884fae/2609.00958)\]  
**\[v1\]** Tue, 1 Sep 2026 09:18:11 UTC (277 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.00958) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
