MEGA Hub

Cross-Domain Hybrid OPD for Generalizable Search Agents

Authors

Do you know Hongzhan Chen?You can claim authorship or link another user.Do you know Xiaoyu Liu?You can claim authorship or link another user.Do you know Dengming Zhang?You can claim authorship or link another user.Do you know Minzhou Huang?You can claim authorship or link another user.Do you know Dongliang Xu?You can claim authorship or link another user.Do you know Jingcheng Xie?You can claim authorship or link another user.Do you know Dongxiang Fang?You can claim authorship or link another user.Do you know Bowen Qin?You can claim authorship or link another user.Do you know Minsheng Hao?You can claim authorship or link another user.Do you know Yaozong Shen?You can claim authorship or link another user.Do you know Xiaojun Quan?You can claim authorship or link another user.Do you know Mona Zhou?You can claim authorship or link another user.Do you know Haosheng Zou?You can claim authorship or link another user.Do you know Jeff Chen?You can claim authorship or link another user.

Abstract

Recent advances in Reinforcement Learning (RL) have substantially improved the capabilities of autonomous search agents, enabling sophisticated planning, and iterative retrieval over dynamic information sources. However, optimizing language models for specialized search behaviors often incurs an alignment tax, where gains in search performance come at the expense of general-purpose capabilities, limiting their effectiveness as universal assistants. In this technical report, we present the training framework behind the Yuanbao search agent, designed to achieve search specialization without sacrificing general intelligence. Built upon the Hunyuan3 architecture, our framework combines agentic reinforcement learning for autonomous search with a cross-domain expert On-Policy Distillation (OPD) pipeline. Experts specializing in complementary general-purpose domains are distilled into the search-specialized student, restoring and further enhancing its broad capabilities. Rather than treating specialization and general capability as competing objectives, our hybrid training strategy jointly optimizes both, effectively mitigating the alignment tax. Extensive experiments demonstrate that the resulting model achieves competitive search performance while consistently improving its general-purpose capabilities, providing a favorable balance between specialized execution and broad generalization in real-world search scenarios.

Community

00