MEGA Hub

MaLViL: Multi-axis Low-rank Vision-LSTM for Medical Image Segmentation

Authors

Do you know Afshin Bozorgpour?You can claim authorship or link another user.Do you know Sina Ghorbani Kolahi?You can claim authorship or link another user.Do you know Moein Heidari?You can claim authorship or link another user.Do you know Ilker Hacihaliloglu?You can claim authorship or link another user.Do you know Dorit Merhof?You can claim authorship or link another user.

Abstract

Vision-LSTM (ViL) enables efficient global modeling, but its cost still scales with the number of spatial tokens, so existing segmenters confine ViL to a coarse bottleneck and lose fine anatomical detail. Rasterizing 2D features into a 1D sequence further breaks adjacency across the orthogonal scan axis. We propose MaLViL, a Multi-axis Low-rank Vision-LSTM network that extends ViL across decoder resolutions. Bidirectional low-rank ViL (Bi-LRViL) reasons on a compact orthonormal subspace and preserves detail through an orthogonal residual; scale-aware SaLViL restores cross-axis neighbors before serialization; and a Cross-Directional Mixer (CDM) fuses orthogonal horizontal and vertical traversal paths. Statistics-Guided Skip Modulation (SGSM) further retains boundary cues in encoder skips. On skin-lesion, ultrasound, and multi-organ CT benchmarks, MaLViL achieves competitive or state-of-the-art segmentation accuracy, while reducing ViL operator memory by up to $83\times$ at fine decoder resolutions. Code is available at: https://github.com/xmindflow/malvil.

Community

00

Publication notes

Author note
Accepted at the MICCAI Workshop on Machine Learning in Medical Imaging (MLMI), 2026