MEGA Hub

An Actionable Diagnosis of Multilingual, Multi-Agent Planning Failures

Authors

Do you know Vikas Pahuja?You can claim authorship or link another user.Do you know Jonathan Brokman?You can claim authorship or link another user.Do you know Omer Hofman?You can claim authorship or link another user.Do you know Tamir Nizri?You can claim authorship or link another user.Do you know Daniel Vishna?You can claim authorship or link another user.Do you know Seraphina Goldfarb-Tarrant?You can claim authorship or link another user.Do you know Kelly Marchisio?You can claim authorship or link another user.Do you know Hisashi Kojima?You can claim authorship or link another user.Do you know Roman Vainshtein?You can claim authorship or link another user.

Abstract

Multilingual multi-agent systems exhibit substantial degradation beyond English, yet prior work rarely identifies how task-critical information is lost when user requests are converted into executable plans. We study the planner in a multi-agent system as the request-to-action interface and derive an actionable taxonomy of planning-grounding failures from failed real-world task executions. LLM-based analysis shows that these failures constitute an increasing share of unsuccessful executions as language-resource availability declines, with the strongest effects in low-resource languages. To test whether the taxonomy supports mitigation, we introduce TART, Taxonomy-Guided Actionable Representation, that makes the taxonomy's key aspects explicit to the planner and downstream sub-agents. Across multiple languages, three LLM backbones, two datasets, and two agentic configurations, TART consistently improves performance. On multilingual GAIA, it raises a state-of-the-art system's accuracy by 5.6 percentage points averaged across eleven languages spanning low- to high-resource settings.

Community

00

Publication notes

Author note
22 pages, 11 figures