MEGA Hub

Towards Hierarchical Structure Understanding of Newspaper Images

Authors

Do you know William Mocaër?You can claim authorship or link another user.Do you know Solène Tarride?You can claim authorship or link another user.Do you know Thomas Constum?You can claim authorship or link another user.Do you know Merveilles Agbeti-Messan?You can claim authorship or link another user.Do you know Tom Simon?You can claim authorship or link another user.Do you know Clément Chatelain?You can claim authorship or link another user.Do you know Stéphane Nicolas?You can claim authorship or link another user.Do you know Pierrick Tranouez?You can claim authorship or link another user.Do you know Sébastien Cretin?You can claim authorship or link another user.Do you know Thierry Paquet?You can claim authorship or link another user.

Abstract

Understanding newspaper images remains a challenging task due to their complex, nested hierarchical structures and dense, heterogeneous layouts. In this paper, we explore two complementary approaches for newspaper structure understanding. First, we present a modular bottom-up pipeline that combines state-of-the-art open-source models: YOLO for layout detection, LayoutReader for reading order prediction, and a custom algorithm for article segmentation. This approach leverages existing robust components while maintaining flexibility and interpretability. Second, we introduce Tiramisu (Tiered Transformers for Hierarchical Structure Understanding), a novel end-to-end transformer-based architecture that explicitly models document hierarchy through an iterative tiered process. Tiramisu performs section and article separation, block localization, semantic categorization, and reading order prediction using highly parallelized attention mechanisms. Finally, we release Finlam La Liberté, a new dataset designed specifically for evaluating hierarchical information retrieval in historical newspapers. Experimental results demonstrate the effectiveness of both approaches in reconstructing complex newspaper hierarchies, with comparative analysis highlighting their respective strengths for scalable document digitization. The Tiramisu training code, including the synthetic newspaper generator, is available at https://git.litislab.fr/tiramisu/tiramisu-newspaper-articles-extractor.

Community

00

Publication notes

Author note
Accepted at ICDAR 2026 Workshop on Historical Document Imaging and Processing (HIP)