MEGA Hub

ContextWeave: A Real-World Workflow Benchmark

Authors

Do you know Bo Wang?You can claim authorship or link another user.Do you know Yuqian Yao?You can claim authorship or link another user.Do you know Enxi Wang?You can claim authorship or link another user.Do you know Luozhijie Jin?You can claim authorship or link another user.Do you know Yang Liu?You can claim authorship or link another user.Do you know Yiran Suo?You can claim authorship or link another user.Do you know Yuxuan Cai?You can claim authorship or link another user.Do you know Enyu Zhou?You can claim authorship or link another user.Do you know Yufei Gao?You can claim authorship or link another user.Do you know Honglin Guo?You can claim authorship or link another user.Do you know Tianyu Huai?You can claim authorship or link another user.Do you know Li Ji?You can claim authorship or link another user.Do you know Zhikai Lei?You can claim authorship or link another user.Do you know Bufan Li?You can claim authorship or link another user.Do you know Lizhi Lin?You can claim authorship or link another user.Do you know Jinxiu Liu?You can claim authorship or link another user.Do you know Jie Yang?You can claim authorship or link another user.Do you know Jiazheng Zhou?You can claim authorship or link another user.Do you know Maosen Zhou?You can claim authorship or link another user.Do you know Pengfang Qian?You can claim authorship or link another user.Do you know Shichun Liu?You can claim authorship or link another user.Do you know Guanshan Liu?You can claim authorship or link another user.Do you know Hao Zheng?You can claim authorship or link another user.Do you know Yunhao Yu?You can claim authorship or link another user.Do you know Hang Yan?You can claim authorship or link another user.Do you know Jihua Kang?You can claim authorship or link another user.Do you know Xinchi Chen?You can claim authorship or link another user.Do you know Xipeng Qiu?You can claim authorship or link another user.

Abstract

Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We introduce ContextWeave, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams. ContextWeave reconstructs privacy-preserved, multi-month workflows of 14 participants into 1,005 executable tasks, including 568 core evaluation tasks, with instructions, containerized environments, trajectories, and task-specific rubrics. It measures workspace quality and alignment with participant-specific preferences, complemented by diagnostics of relevance, continuity, solvability, and robustness to misleading recall. Across six memory components under a fixed model, the strongest configuration raises Workspace Score from 68.08 to 78.20 and Preference Score from 41.50 to 70.60. With a fixed memory component, recall improves both outcomes for all five tested base models, although gains vary substantially. Our analysis shows that actionable, experience-rich memory supports workflow continuation and reduces redundant exploration more effectively than compact summaries, while it can also be more susceptible to misleading recall. These findings motivate memory systems that optimize not only retrieval relevance but also reliable use during execution.

Community

00