MEGA Hub

An end-to-end-trained vision-language model for native-language prostate pathology report generation

Authors

Do you know Christian Grashei?You can claim authorship or link another user.Do you know Fabian Gülhan?You can claim authorship or link another user.Do you know Maximilian Legnar?You can claim authorship or link another user.Do you know Fabian Stögbauer?You can claim authorship or link another user.Do you know Cleo-Aron Weis?You can claim authorship or link another user.Do you know Carolin Mogler?You can claim authorship or link another user.Do you know Peter Schüffler?You can claim authorship or link another user.

Abstract

Prostate cancer is among the most frequently diagnosed malignancies worldwide, and structured reporting of each biopsy core burdens pathologists. Existing tools frame this as classification, leaving pathologists to assemble coherent reports, while many slide-level vision-language models rely on English-centric encoders that transfer poorly to other clinical languages. We present a slide-level framework generating prostate biopsy reports that is language-independent by construction: tokenizer and model are trained from scratch, demonstrated here in German. To address paired-data scarcity, an automated pipeline uses a locally deployed large language model to split composite reports into core-specific image-text pairs, yielding 17,344 pairs from 2,402 historical cases without manual annotation. Evaluated for clinical attributes rather than linguistic similarity, the model achieves 96.2% F1 for malignancy detection and 65.2% for Gleason grading, competitive with an FDA-cleared classifier. Grading is further validated on three external cohorts with latent-space augmentation. Institutions can thus train native-language reporting models on their own archives.

Community

00