MEGA Hub

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

Authors

Do you know Tencent WorkBuddy Bench Team?You can claim authorship or link another user.Do you know Siqi Cai?You can claim authorship or link another user.Do you know Shaopeng Chen?You can claim authorship or link another user.Do you know Xiang Fei?You can claim authorship or link another user.Do you know Yong Mao?You can claim authorship or link another user.Do you know Zihan Xu?You can claim authorship or link another user.Do you know Zhiheng Lyu?You can claim authorship or link another user.Do you know Zhijian Shao?You can claim authorship or link another user.Do you know Yuchen Shi?You can claim authorship or link another user.Do you know Shuwen Zhang?You can claim authorship or link another user.Do you know Chaofan Qiu?You can claim authorship or link another user.Do you know Linjie Che?You can claim authorship or link another user.Do you know Xiaoxi Zhao?You can claim authorship or link another user.Do you know Feng Wu?You can claim authorship or link another user.Do you know Kai Zhang?You can claim authorship or link another user.Do you know Chaofan Zhu?You can claim authorship or link another user.Do you know Yubin Qi?You can claim authorship or link another user.Do you know Xiaoyun Liang?You can claim authorship or link another user.Do you know Peijie Dong?You can claim authorship or link another user.Do you know Yunhao Zhang?You can claim authorship or link another user.Do you know Yuanjie Zhu?You can claim authorship or link another user.Do you know Ling Jiang?You can claim authorship or link another user.Do you know Xianjun Zhang?You can claim authorship or link another user.Do you know Zhehang Chu?You can claim authorship or link another user.Do you know Anyuan Sang?You can claim authorship or link another user.Do you know Zhen Feng?You can claim authorship or link another user.Do you know Sen Nie?You can claim authorship or link another user.Do you know Shi Wu?You can claim authorship or link another user.Do you know Yuanzhen Xu?You can claim authorship or link another user.Do you know Xin Li?You can claim authorship or link another user.Do you know Ning Yang?You can claim authorship or link another user.Do you know Zhiqiang Dong?You can claim authorship or link another user.Do you know Hande Dong?You can claim authorship or link another user.Do you know Qiang Lin?You can claim authorship or link another user.Do you know Yi Liu?You can claim authorship or link another user.Do you know Yunsheng Wu?You can claim authorship or link another user.Do you know Ke Li?You can claim authorship or link another user.Do you know Xing Sun?You can claim authorship or link another user.

Abstract

We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation framework for constructing and running distribution-informed coding-agent tasks across four work domains - Code, Web, Office, and Security. Rather than adapting public issue text, every task is reverse-engineered from a real commit, pull request, or business scenario and rewritten as a short, colloquial, role-played request, so that a task's prompt is not recoverable by web-searching the underlying issue, pull request, or commit thread. Because the dataset is released openly - task directories, environment images, evaluation harness, tests, and reference solutions - contamination resistance rests on this construction together with dataset versioning rather than on secrecy. The four subsets - repository-level engineering, front-end development, office and business workflows, and red-/blue-team security - probe complementary facets of real work, each with its own verification style. All are packaged in a uniform task-directory format and run, under a uniform and reproducible protocol, on two agent harnesses (CodeBuddy Code and Claude Code); the full open release makes the benchmark reproducible end to end and directly auditable, since any third party can re-run each task and inspect its content. Because each subset uses a different scoring instrument, scores are not comparable across subsets and the suite reports no suite-wide average. We report a cross-model leaderboard across several model families.

Community

00

Publication notes

Author note
30 pages, 9 figures. Project page: https://workbuddybench.com/ ; code: https://github.com/Tencent/workbuddy-bench ; dataset: https://huggingface.co/datasets/tencent/workbuddy-bench