MEGA Hub

Step-Level Preference Learning for Generative Agents in Social Simulations

Authors

Do you know Wenchang Gao?You can claim authorship or link another user.Do you know Pingyue Sheng?You can claim authorship or link another user.Do you know Lanlan Qiu?You can claim authorship or link another user.Do you know Yunfei Ma?You can claim authorship or link another user.Do you know Jian Zhao?You can claim authorship or link another user.Do you know Baicheng Chen?You can claim authorship or link another user.Do you know Kangda Wang?You can claim authorship or link another user.Do you know Yuyang Tian?You can claim authorship or link another user.Do you know Shunqiang Mao?You can claim authorship or link another user.Do you know Tianxing He?You can claim authorship or link another user.

Abstract

Large language model (LLM)-based generative agents simulate human behavior through long-horizon decision-making processes that comprise intermediate steps such as planning, memory retrieval, reflection, and action selection. However, fine-grained human annotations of these intermediate steps remain scarce, and existing agents are not grounded in human preferences over such intermediate decisions. To address this gap, we introduce \method, an interactive simulation interface that enables us to collect step-level human preference supervision over agent decision trajectories, leading to a dataset of 57K fine-grained annotations. We conduct step-level preference learning on open-weight language models using supervised finetuning and direct preference optimization on this data, consistently improving simulation fidelity, coordination, and interaction quality, and inducing more socially effective agent behavior. Our results show that step-level human supervision is an effective training signal for improving both local decision quality and long-horizon agent behavior.

Community

00

Publication notes

Author note
WAICA2026