MEGA Hub

CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations

Authors

Do you know Xingyu Yan?You can claim authorship or link another user.Do you know Tingting Dai?You can claim authorship or link another user.Do you know Antonio De Domenico?You can claim authorship or link another user.Do you know Mohamed Sana?You can claim authorship or link another user.Do you know Nicola Piovesan?You can claim authorship or link another user.Do you know Changchang Li?You can claim authorship or link another user.Do you know Bowen Liu?You can claim authorship or link another user.Do you know Kun Jiang?You can claim authorship or link another user.Do you know Mengjie Zhang?You can claim authorship or link another user.Do you know Dingcheng Shan?You can claim authorship or link another user.Do you know Jing-Cheng Pang?You can claim authorship or link another user.Do you know Chenwei Wu?You can claim authorship or link another user.Do you know Sijie Wu?You can claim authorship or link another user.Do you know Lianying Chao?You can claim authorship or link another user.Do you know Haoran Cai?You can claim authorship or link another user.Do you know Jiantao Ye?You can claim authorship or link another user.Do you know Xubin Li?You can claim authorship or link another user.Do you know Simon Mark Lucas?You can claim authorship or link another user.Do you know Xin Chen?You can claim authorship or link another user.

Abstract

Agents are increasingly considered for automating network operations and maintenance, where engineers must diagnose network faults, optimize configurations to enhance services, and reduce operational costs while acting under strict constraints. However, existing evaluations fail to accurately model real network characteristics or assess agents under partially observable telecom environments with diverse vendors, devices, protocols, and interfaces. In this paper, we introduce CTBench, a public benchmark for assessing whether an agent behaves like a competent telecom troubleshooting engineer. CTBench focuses on root cause analysis and path restoration. Each task is constructed by experts and annotated with rich task metadata, including golden evidence steps. CTBench uses expert-grounded metrics that evaluate both final answers and the diagnostic evidence. Experiments with representative harness-model combinations show that state-of-the-art agents perform very well at identifying endpoints in path-restoration tasks but, more generally, underperform in root cause analysis. In particular, agents struggle with interface state, link-layer, service-management, and other operational faults. Most importantly, even when agents produce plausible or correct final answers, they often fail to provide the evidence-grounded diagnoses required in operational practice. Our results further show that path restoration is generally more resource expensive, yet larger resource usage does not necessarily translate into better diagnosis.

Community

00