MEGA Hub

MELLON - Multimodal Enhanced LLM for Online Navigation

Authors

Do you know Ruiyu Li?You can claim authorship or link another user.Do you know Haoyang Cai?You can claim authorship or link another user.Do you know Zhitong Guo?You can claim authorship or link another user.Do you know Tong Hu?You can claim authorship or link another user.

Abstract

Web navigation agents are capable of addressing various types of tasks on different websites. Current baselines on web navigation are either unimodal or lack strong reasoning abilities given multimodal inputs. Focusing on the WebShop benchmark, a real-world website simulation, we explore the alignment of text and images, as well as multimodal reasoning and planning abilities, to enhance the performance of web navigation agents. We propose three innovative multimodal enhancements: Multimodal Enhanced LLM for Online Navigation (MELLON), VQAgent, and Multimodal Ranker. MELLON demonstrates a significant improvement in task completion accuracy, with a 9.26% increase after just one epoch of training. Our findings suggest the necessity of further exploration into multimodal approaches, with a focus on more extensive training and alignment strategies to enhance the effectiveness of web navigation agents.

Community

00