MEGA Hub

Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?

Authors

Do you know Yuyang Dai?You can claim authorship or link another user.Do you know Xueqing Peng?You can claim authorship or link another user.Do you know Yuxia Wang?You can claim authorship or link another user.Do you know Preslav Nakov?You can claim authorship or link another user.Do you know Zhuohan Xie?You can claim authorship or link another user.

Abstract

Large language models are increasingly applied as autonomous decision-making agents. However, in executive business decisions, existing benchmarks are limited to textonly settings. This makes it unclear whether models can perceive visual business evidence and effectively integrate it to improve decision quality. We introduce C-SUITEBENCH, a controlled multimodal benchmark that includes five decision tasks under paired text-only and multimodal conditions across 50 scenarios. We place nine frontier models in the role of a chief executive officer and evaluate their decision-making ability. Multimodal inputs consistently improve evidence-centric reasoning, with the largest and most reliable gains appearing in risk forecasting and board-facing justification. However, we uncover a multimodal integration paradox: adding visual business information degrades constrained resource allocation for all nine models, even as visual grounding itself improves. Ablation experiments reveal that this failure emerges from signal crowding, although each visual channel helps individually, their combination disrupts constraint satisfaction during decoding. These findings demonstrate that visual perception and constrained action are separable bottlenecks in multimodal agents, and that indiscriminate visual augmentation can harm high-stakes decision making, motivating selective grounding strategies for future executive AI systems.

Community

00

Publication notes

Author note
25 pages