MEGA Hub

Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

Authors

Do you know Uma Ranjan?You can claim authorship or link another user.Do you know Kunal Tilaganji?You can claim authorship or link another user.Do you know Aditya Koul?You can claim authorship or link another user.Do you know Anurag Mahipal?You can claim authorship or link another user.Do you know Dashpreet Singh?You can claim authorship or link another user.Do you know Hriday Rana?You can claim authorship or link another user.Do you know Manan Jain?You can claim authorship or link another user.Do you know Sidharth Gupta?You can claim authorship or link another user.Do you know Ajo Babu George?You can claim authorship or link another user.Do you know Vineeth Balasubramanian?You can claim authorship or link another user.Do you know Nagarajan Natarajan?You can claim authorship or link another user.Do you know Amit Sharma?You can claim authorship or link another user.

Abstract

Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models to abstain when uncertain improves reliability but introduces a coverage accuracy tradeoff. We propose a two-stage framework for medical hypothesis verification in multiple-choice settings that manages this tradeoff through targeted ontology grounding, applied only when the model abstains. We show that abstention is not random but reflects genuine uncertainty, with abstained predictions associated with lower confidence. Across two frontier models (GPT-5.5, accessed via the Azure OpenAI API, and DeepSeek-R1), the proposed framework improves question-level accuracy by 9.6 percentage points (82.9% to 92.5%) and hypothesis-level accuracy by 4.2 percentage points (92.0% to 96.2%). Our experiments conducted on MedReason and MedQA show that abstention can be repurposed as a control signal for selective reasoning refinement, achieving knowledge-graph-level performance without explicit knowledge graph construction.

Community

00

Publication notes

Author note
Findings Track at the Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)