MEGA Hub

OntoBook: Ontology-Grounded Synthetic Textbooks for Medical Encoder Pretraining

Authors

Do you know Rian Touchent?You can claim authorship or link another user.Do you know Éric de la Clergerie?You can claim authorship or link another user.

Abstract

We present OntoBook, a method that converts medical ontology structure into pretraining signal for encoder language models. Our approach has three stages: random walks through ontology graphs capture hierarchical and causal relations between medical codes, a large language model reformulates these walks into fluent textbook-style prose, and the resulting text is used to train ModernCamemBERT, a 149M-parameter French encoder, with two objectives on the same data: masked language modeling and relation prediction between code pairs. On three French medical coding benchmarks (FRACCO, Cantemist-FR, Distemist-FR), OntoBook achieves significant improvements over MLM-only pretraining, with +2.5 micro-F1 on FRACCO and +8.0 micro-F1 on Distemist. We find that alignment between objectives is necessary: misaligned training, where each task uses different data, causes a 30-point degradation. We release 1.3 million LLM-reformulated medical textbooks across three French ontologies (CIM-10, CCAM, ATC) and pretrained model checkpoints.

Community

00

Publication notes

Journal
Proceedings of Knowledge Graphs and Large Language Models Workshop, May 2026, Palma de Mallorca, Spain