MEGA Hub

Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs

Authors

Do you know Vu Duc Anh?You can claim authorship or link another user.Do you know Nhat M. Hoang?You can claim authorship or link another user.Do you know Do Xuan Long?You can claim authorship or link another user.Do you know Cong-Duy Nguyen?You can claim authorship or link another user.Do you know Ponhvoan Srey?You can claim authorship or link another user.Do you know Luu Anh Tuan?You can claim authorship or link another user.

Abstract

Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language models (LLMs). In this work, we propose Self-Fix Step-DPO (SFS-DPO), a reinforcement learning based, two-stage framework for step-level self-verification and self-correction. The first stage strengthens step-level reasoning via step-level preference optimization, while the second stage explicitly trains models to self-verify and self-correct. We further introduce a teacher-assisted variant, SFS-DPO-R, which incorporates explanatory rationales for error verification to provide stronger corrective signals. Comprehensive in-domain and out-of-domain evaluations across multiple LLMs demonstrate that SFS-DPO and SFS-DPO-R consistently outperform prior step-level training baselines. Our analysis further reveals improvements in self-correction frequency and effectiveness, highlighting the importance of strengthening step-level reasoning for robust performance.

Community

00