MEGA Hub

Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning

Authors

Do you know Yongqi Tong?You can claim authorship or link another user.Do you know Zhenyu Zhang?You can claim authorship or link another user.Do you know Zimi Liu?You can claim authorship or link another user.Do you know Kewei Fu?You can claim authorship or link another user.Do you know Mingli Song?You can claim authorship or link another user.Do you know Haofei Zhang?You can claim authorship or link another user.Do you know Junshao Zhang?You can claim authorship or link another user.Do you know Hong Zhu?You can claim authorship or link another user.Do you know Jiang-Ming Yang?You can claim authorship or link another user.Do you know Xin Zhang?You can claim authorship or link another user.Do you know Jianshe Li?You can claim authorship or link another user.

Abstract

Answer-only reinforcement learning (RL) trains reasoning models to solve fully specified problems, but many realistic queries omit a premise needed for a unique answer. In this setting, the useful response is not always refusal: the model should ask for the missing premise, condition its answer on the unknown quantity, or abstain when no informative conditional response is available. We present \emph{Ask-Condition-Abstain Reinforcement Learning} (ACA-RL), a data-augmented RL framework for this setting. Its reasoning-graph-guided pipeline converts well-posed problems into missing-premise training instances with localized gap annotations; ACA-RL then trains on these instances with a structured reward over five observable response behaviors. We also introduce the \emph{Missing-Premise Benchmark} (MPB), a 274-instance human-verified benchmark spanning mathematical, logical, and real-world word problems. Across Qwen3 and Llama models, ACA-RL consistently improves on MPB while preserving competitive performance on well-posed reasoning tasks. Together with the released code, MPB, and training data, this work supports a new mission for NLP evaluation: measuring whether models can recognize when a task is underdetermined and handle uncertainty, not only whether they can answer fully specified questions.

Community

00