MEGA Hub

ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons

Authors

Do you know Adrian Li?You can claim authorship or link another user.Do you know Kelong Mao?You can claim authorship or link another user.Do you know Yudong Guo?You can claim authorship or link another user.Do you know Heming Xia?You can claim authorship or link another user.Do you know Xinwei Yang?You can claim authorship or link another user.Do you know Lirui Luo?You can claim authorship or link another user.Do you know Jace Wong?You can claim authorship or link another user.Do you know Pu Yao?You can claim authorship or link another user.Do you know Sulong Xu?You can claim authorship or link another user.Do you know Simiu Gu?You can claim authorship or link another user.

Abstract

Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shopping tasks arise in device setup, meal preparation, event planning, and group takeout ordering, requiring joint reasoning about item compatibility, availability, store-level requirements, delivery fees, coupons, and budgets. Evaluation is challenging because multiple baskets may satisfy the same request, making exact-match metrics unsuitable, whereas semantic evaluation alone cannot detect infeasible orders, invalid coupon combinations, or incorrect payments. We introduce ComboShoppingBench, an agentic shopping benchmark for open-ended yet verifiable basket construction in a simulated commerce and takeout environment. During task synthesis, an exploration agent constructs a feasible and semantically coherent basket of purchasable products; this witness guides the generation of coupons, budget constraints, user queries, and aligned evaluation rubrics. During evaluation, LLM judges assess semantic satisfaction, response quality, and claim faithfulness, while deterministic validation checks product-ID validity, budget compliance, and coupon optimality. Experiments with diverse LLM agents demonstrate that even strong agents struggle on ComboShoppingBench, highlighting substantial room for improvement in reliable, constraint-aware combo shopping.

Community

00