MEGA Hub

When Text and Numbers Disagree: Evidence Arbitration in Large Language Models

Authors

Do you know Mattia Carletti?You can claim authorship or link another user.Do you know Edward Phillips?You can claim authorship or link another user.Do you know Fredrik K. Gustafsson?You can claim authorship or link another user.Do you know Patitapaban Palo?You can claim authorship or link another user.Do you know Lei Clifton?You can claim authorship or link another user.Do you know Danielle Belgrave?You can claim authorship or link another user.Do you know Xiao Gu?You can claim authorship or link another user.Do you know David A. Clifton?You can claim authorship or link another user.

Abstract

Large language models (LLMs) are increasingly used in settings where textual summaries, numerical observations, and external tool outputs may provide conflicting evidence. We study how LLMs arbitrate between such sources when they support opposing decisions. To do so, we introduce a controlled synthetic benchmark in which latent risk trajectories generate both numerical time series and natural language summaries, allowing us to construct conflicts where exactly one evidence source is aligned with the ground-truth label. This design lets us independently manipulate modality, temporal recency, source reliability, and evidence provenance. Across open-weight instruction-tuned models, we find that arbitration behaviour is systematic rather than random: models exhibit distinct text-versus-number preferences, follow temporal recency more consistently than explicit reliability cues, and can over-rely on external forecasts even when they conflict with direct contextual evidence. These results suggest that current LLMs often rely on heuristic arbitration strategies when integrating heterogeneous evidence, highlighting a failure mode for tool-augmented decision systems.

Community

00