---
格式版本: 2
标题: "CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations"
原文链接: "https://arxiv.org/abs/2608.12002"
发布日期: "2026-08-12"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 12 Aug 2026 12:37:02 UTC (503 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-13T18:33:52+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-13T18:24:32+08:00"
入库时间: "2026-08-13T10:33:52.662Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=AI&searchtype=all"
匹配关键词:
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI质检状态: "评分失败"
AI评分尝试次数: 1
AI评分错误类型: "service_error"
AI评分错误: "LLM call failed; tried model chain: ali-deepseek-v4-flash -> tx-deepseek-v4-flash | Model ali-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131033552662957678268d9d6j4RfaRKy)\",\"type\":\"new_api_error\"}} | Model tx-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131033558597582508268d9d6TlS3pEHZ)\",\"type\":\"new_api_error\"}}"
AI评分开始时间: "2026-08-13T10:33:52.670Z"
AI评分结束时间: "2026-08-13T10:33:55.994Z"
AI摘要: "论文提出CTBench，一个用于评估AI智能体在真实电信网络运维中故障排查能力的公开基准，聚焦根因分析和路径恢复。实验显示，现有最先进智能体在路径恢复端点识别表现良好，但根因分析普遍不佳，即使给出看似正确的最终答案，也常缺乏基于证据的诊断。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T03:40:15.220Z"
采集批次: "2026年8月13日18点24分27秒"
采集批次ID: "20260813-182427-1564e572"
去重键: "https://arxiv.org/abs/2608.12002"
---

## Computer Science > Artificial Intelligence

## Title:CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations

Authors:[Xingyu Yan](https://arxiv.org/search/cs?searchtype=author&query=Yan,+X), [Tingting Dai](https://arxiv.org/search/cs?searchtype=author&query=Dai,+T), [Antonio De Domenico](https://arxiv.org/search/cs?searchtype=author&query=De+Domenico,+A), [Mohamed Sana](https://arxiv.org/search/cs?searchtype=author&query=Sana,+M), [Nicola Piovesan](https://arxiv.org/search/cs?searchtype=author&query=Piovesan,+N), [Changchang Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+C), [Bowen Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+B), [Kun Jiang](https://arxiv.org/search/cs?searchtype=author&query=Jiang,+K), [Mengjie Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+M), [Dingcheng Shan](https://arxiv.org/search/cs?searchtype=author&query=Shan,+D), [Jing-Cheng Pang](https://arxiv.org/search/cs?searchtype=author&query=Pang,+J), [Chenwei Wu](https://arxiv.org/search/cs?searchtype=author&query=Wu,+C), [Sijie Wu](https://arxiv.org/search/cs?searchtype=author&query=Wu,+S), [Lianying Chao](https://arxiv.org/search/cs?searchtype=author&query=Chao,+L), [Haoran Cai](https://arxiv.org/search/cs?searchtype=author&query=Cai,+H), [Jiantao Ye](https://arxiv.org/search/cs?searchtype=author&query=Ye,+J), [Xubin Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+X), [Simon Mark Lucas](https://arxiv.org/search/cs?searchtype=author&query=Lucas,+S+M), [Xin Chen](https://arxiv.org/search/cs?searchtype=author&query=Chen,+X)

[View PDF](https://arxiv.org/pdf/2608.12002) [HTML (experimental)](https://arxiv.org/html/2608.12002v1)

> Abstract:Agents are increasingly considered for automating network operations and maintenance, where engineers must diagnose network faults, optimize configurations to enhance services, and reduce operational costs while acting under strict constraints. However, existing evaluations fail to accurately model real network characteristics or assess agents under partially observable telecom environments with diverse vendors, devices, protocols, and interfaces. In this paper, we introduce CTBench, a public benchmark for assessing whether an agent behaves like a competent telecom troubleshooting engineer. CTBench focuses on root cause analysis and path restoration. Each task is constructed by experts and annotated with rich task metadata, including golden evidence steps. CTBench uses expert-grounded metrics that evaluate both final answers and the diagnostic evidence. Experiments with representative harness-model combinations show that state-of-the-art agents perform very well at identifying endpoints in path-restoration tasks but, more generally, underperform in root cause analysis. In particular, agents struggle with interface state, link-layer, service-management, and other operational faults. Most importantly, even when agents produce plausible or correct final answers, they often fail to provide the evidence-grounded diagnoses required in operational practice. Our results further show that path restoration is generally more resource expensive, yet larger resource usage does not necessarily translate into better diagnosis.

| Subjects: | Artificial Intelligence (cs.AI) |
| --- | --- |
| Cite as: | [arXiv:2608.12002](https://arxiv.org/abs/2608.12002) \[cs.AI\] |
|  | (or [arXiv:2608.12002v1](https://arxiv.org/abs/2608.12002v1) \[cs.AI\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.12002](https://doi.org/10.48550/arXiv.2608.12002) |

## Submission history

From: Mohamed Sana \[[view email](https://arxiv.org/show-email/8772eedc/2608.12002)\]  
**\[v1\]** Wed, 12 Aug 2026 12:37:02 UTC (503 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.12002) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
