MEGA Hub

Robust Interpretation of Historical Documents in Knowledge Graphs Through Query Inference and Execution

Authors

Do you know Sebastià Nicolau?You can claim authorship or link another user.Do you know Adrià Molina?You can claim authorship or link another user.Do you know Oriol Ramos Terrades?You can claim authorship or link another user.Do you know Josep Lladós?You can claim authorship or link another user.

Abstract

The emergence of Large Language Models (LLMs) has redefined how users interact with information in digital environments. However, their widespread and often indiscriminate integration has raised significant concerns regarding reliability and trustworthiness issues that are particularly critical when accessing digital libraries and historical archives. How can one leverage the generalization capacity of an LLM without losing the level of accountability required for an archival institution? In this paper, we present an agentic retrieval system designed to deliver more accurate and verifiable access to historical data while preserving much of the flexibility associated with unconstrained LLMs. As a contribution to historical document analysis, we compare traditional Retrieval-Augmented Generation (RAG) with an agentic GraphRAG architecture in their ability to deliver historical information under realistic conditions, including the presence of OCR and transcription errors. We introduce a semi-symbolic framework that integrates word-spotting techniques for post-OCR correction with a knowledge graph representation that enables the agent to access information through synthesized queries. The interleaved collaboration between word spotting and code generation allows the agent to construct strong retrieval queries that are robust to misinterpretation and hallucination, while still leveraging approximate search when noise and uncertainty, common in historical document analysis, would otherwise hinder precise retrieval.

Community

00

Publication notes

Author note
Accepted at ICDAR2026