MEGA Hub

MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models

Authors

Do you know Yuan Wang?You can claim authorship or link another user.Do you know Hualiang Wang?You can claim authorship or link another user.Do you know Yixin Chen?You can claim authorship or link another user.Do you know Songtao Jiang?You can claim authorship or link another user.Do you know Shujian Gao?You can claim authorship or link another user.Do you know Jiaming Lin?You can claim authorship or link another user.Do you know Siming Fu?You can claim authorship or link another user.Do you know Jian Wu?You can claim authorship or link another user.Do you know Zuozhu Liu?You can claim authorship or link another user.

Abstract

Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain challenging. Existing approaches either verbalize regions as coordinate strings or rely on external modules that decouple perception from understanding, creating representation gaps for region-language alignment. We present MedUP, a Med-VLM that natively unifies perception and understanding within a shared token space. At its core lies UniMedTok, a region tokenizer that encodes masks as discrete tokens in the LLM vocabulary, enabling the model to seamlessly interleave mask tokens with text. We curate UniMed-Train, a 1.84M-instance corpus spanning text-guided segmentation, region-grounded understanding, medical VQA and CoT-based segmentation, and introduce UniMed-Bench for unified evaluation. Extensive experiments show that MedUP outperforms native, agentic, and dual-decoder Med-VLMs across all tasks while remaining competitive with specialist segmentors, demonstrating the strong potential of unified understanding and perception modeling.

Community

00

Publication notes

Author note
10 pages, 6 figures