MEGA Hub

Semantically Similar, Logically Distinct: Diagnosing the Semantic-Answerability Gap in Table RAG

Authors

Do you know Jiaming Tian?You can claim authorship or link another user.Do you know Liyao Li?You can claim authorship or link another user.Do you know Wentao Ye?You can claim authorship or link another user.Do you know Haobo Wang?You can claim authorship or link another user.Do you know Lihua Yu?You can claim authorship or link another user.Do you know Zujie Ren?You can claim authorship or link another user.Do you know Gang Chen?You can claim authorship or link another user.Do you know Junbo Zhao?You can claim authorship or link another user.

Abstract

Tables are a critical knowledge source in retrieval-augmented generation (RAG), but a retrieved table may lack sufficient evidence to answer a query, a property we call answerability. While answerability broadly concerns whether a source or collection of sources contains sufficient evidence, retrieval models optimized for semantic relevance do not guarantee it even in the single-source case, creating a fundamental mismatch. To study this, we introduce TCR-Bench, a diagnostic benchmark for Table Content-level Answerability in RAG, built around sibling tables, i.e., tables with highly similar schemas but subtle content differences. On TCR-Bench, the dense retrievers we evaluate persistently exhibit a Semantic-Answerability Gap: they often retrieve the correct sibling group yet struggle to pinpoint the uniquely answerable table within it, dropping QA performance from 0.755 (oracle) to 0.330 (top-5 retrieved). Our analysis suggests this gap is associated with semantic accumulation, schema-level cue dependence, and weak row-column binding. As a diagnostic probe into the source of this gap, we test whether a lightweight two-stage pipeline, Answerability-Aware Reranking (AAR), applying direct query-table answerability judgment, can recover performance: it raises top-1 target retrieval from 18.2% to 57.4%, and this large gain is itself evidence that much of the observed failure reflects a missing answerability verification step, rather than an inherent limitation of model capacity alone.

Community

00