MEGA Hub

PAMT: Process-Aligned Reinforcement Learning for Multi-Domain Machine Translation

Authors

Do you know Yongshi Ye?You can claim authorship or link another user.Do you know Biao Fu?You can claim authorship or link another user.Do you know Chongxuan Huang?You can claim authorship or link another user.Do you know Yidong Chen?You can claim authorship or link another user.Do you know Xiaodong Shi?You can claim authorship or link another user.

Abstract

Multi-domain machine translation (MDMT) requires more than fluent generation: it demands domain-sensitive translation decisions such as domain disambiguation, terminology control, and stylistic adaptation. Large reasoning models (LRMs) make such decisions explicit through intermediate translation steps, but our analysis across 15 domains and four translation directions shows that this explicit reasoning is double-edged: it improves long-form and high-difficulty translation, yet often drifts in terminology-intensive and stylistically constrained settings. We trace this failure to a credit-assignment bottleneck: existing methods optimize final outputs or coarse trajectories, but cannot identify which translation steps actually help the final translation. To address this, we propose PAMT, a process-aligned training framework that combines cold-start domain-aware Long-CoT supervision with reinforcement learning. PAMT uses sequence-level format and outcome rewards for the final translation, together with a step-level process reward that measures how much each explicit translation step increases the likelihood of the reference translation. Across two backbones, PAMT improves over base models, outperforms MT-specialized baselines on average, and remains competitive with strong LLMs/LRMs across in-domain, OOD, and multilingual settings.

Community

00

Publication notes

Author note
23 pages, 10 figures, and 18 tables