---
格式版本: 2
标题: "TaxCE : A Framework for Automated Taxonomy Construction and Evaluation at Scale"
原文链接: "https://arxiv.org/abs/2608.30614"
发布日期: "2026-08-31"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Mon, 31 Aug 2026 11:25:25 UTC (430 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-02T03:50:33+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-02T03:49:17+08:00"
入库时间: "2026-09-01T19:50:33.686Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=Scale-up&searchtype=all"
匹配关键词:
  - "Scale-up"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 5
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "该资料为arXiv论文TaxCE，主题是NLP自动分类构建与评估，与超节点/AI Rack/机柜级AI基础设施完全无关。仅命中搜索词'Scale-up'，实际为Scale，非技术扩展。"
AI质检模型: "zj-deepseek-v4-flash"
AI质检时间: "2026-09-02T03:51:07+08:00"
AI主题相关性: 0
AI来源权威性: 2
AI新颖性: 3
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 0
AI摘要: "TaxCE提出一种从原始反馈文本自动构建多层级分类法的框架，并配套EEG三项评估指标和指标回环迭代修正机制。实验显示其在排他性、穷尽性和粒度上较最强基线平均提升11.8、20.5和15.7个百分点，人工评估也确认了更优的分类质量与可导航性。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-01T23:06:24.320Z"
采集批次: "2026年9月1日23点56分52秒"
采集批次ID: "20260901-235652-761"
去重键: "https://arxiv.org/abs/2608.30614"
---

## Computer Science > Computation and Language

## Title:TaxCE: A Framework for Automated Taxonomy Construction and Evaluation at Scale

[View PDF](https://arxiv.org/pdf/2608.30614) [HTML (experimental)](https://arxiv.org/html/2608.30614v1)

> Abstract:Organizing unstructured feedback text into hierarchical taxonomy is a fundamental challenge in NLP, particularly in domains where feedback arrives at massive scale in varied forms such as reviews, transcripts, and surveys. Existing approaches either produce shallow hierarchies, neglect long-tail topics, or lack rigorous evaluation frameworks. We present TaxCE, a fully automated framework that constructs multi-level hierarchical taxonomies from raw text through progressive condensation of corpus content into actionable segments, deduplicated semantic units, and granular topics with definitions, which are then organized bottom-up into a hierarchy with corpus-groundedness. We also introduce three corpus-grounded evaluation metrics, Exclusivity, Exhaustivity, and Granularity (EEG), and integrate them into a metrics-in-the-loop iterative refinement mechanism that diagnoses deficiencies and applies targeted corrections until convergence. Extensive experiments demonstrate that TaxCE consistently outperforms existing baselines spanning classical topic models, neural methods, and LLM-based approaches, with average improvements of 11.8, 20.5, and 15.7 percentage points in exclusivity, exhaustivity, and granularity respectively over the strongest baseline. Human evaluation further confirms superior taxonomy quality, actionability, and navigability.

| Comments: |  |
| --- | --- |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | [arXiv:2608.30614](https://arxiv.org/abs/2608.30614) \[cs.CL\] |
|  | (or [arXiv:2608.30614v1](https://arxiv.org/abs/2608.30614v1) \[cs.CL\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.30614](https://doi.org/10.48550/arXiv.2608.30614) |

## Submission history

From: Sandeep Sricharan Mukku \[[view email](https://arxiv.org/show-email/7fe89452/2608.30614)\]  
**\[v1\]** Mon, 31 Aug 2026 11:25:25 UTC (430 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.30614) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
