Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation
Abstract
Open multilingual translation models are improved via group relative policy optimization with reference-free quality rewards and checkpoint interpolation, surpassing strong open and proprietary baselines.
We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO) with a reward that averages two reference-free quality estimation models and is gated by language identification. We then linearly interpolate the supervised fine-tuning (SFT) and reinforcement learning (RL) model checkpoints to obtain MiLMMT-46-v1.0. Across 46 languages, the resulting models consistently improve translation quality over their SFT counterparts, outperform strong recent open baselines, including Seed-X, HY-MT2, and TranslateGemma, and achieve leading reference-free scores against evaluated proprietary systems such as Google Translate, Gemini 3 Pro, and GPT-5. We further investigate on-policy distillation and find that it reaches, but does not surpass, the quality frontier achieved by RL with checkpoint interpolation. We release the models and code to facilitate future research.
Community
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- PAMT: Process-Aligned Reinforcement Learning for Multi-Domain Machine Translation (2026)
- Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation (2026)
- Reasoning Before Translation: Enhancing Legal Machine Translation with Structured Reasoning (2026)
- CAT-Translate: Building Compact Open-Source Models for Japanese-English Translation (2026)
- MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages (2026)
- Distill Where the Student Goes: Teacher-Regularized RL for English-Evidence Cross-Lingual RAG (2026)
- Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.10812 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 3
xiaomi-research/MiLMMT-46-4B-v1.0
Datasets citing this paper 0
No dataset linking this paper