MiLMMT翻譯提質

2026年8月13日 00:00
站內 AI 整理稿

Papers arxiv:2608.

10812 Copy markdown Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation Published on Aug 11 · Submitted by Pengzhi Gao on Aug 12 · Xiaomi Research Upvote 10 +2 Authors: Chris Han ,Pengzhi Gao ,Pei Fu ,Jian Luan Abstract Open multilingual translation models are improved via group relative policy optimization with reference-free quality rewards and checkpoint interpolation, surpassing strong open and proprietary baselines.

Generated by thinkingmachines/Inkling-Small We study reference-free post-training for multilingual machine translation with open large language models.Starting from the supervised-finetuned MiLMMT-46-v0.

1 models, we apply Group Relative Policy Optimization (GRPO) with a reward that averages two reference-free quality estimation models and is gated by language identification.

We then linearly interpolate the supervised fine-tuning (SFT) and reinforcement learning (RL) model checkpoints to obtain MiLMMT-46-v1.0.

Across 46 languages, the resulting models consistently improve translation quality over their SFT counterparts, outperform strong recent open baselines, including Seed-X, HY-MT2, and TranslateGemma, and achieve leading reference-free scores against evaluated proprietary systems such as Google Translate, Gemini 3 Pro, and GPT-5.

We further investigate on-policy distillation and find that it reaches, but does not surpass, the quality frontier achieved by RL with checkpoint interpolation.We release the models and code to facilitate future research.

View arXiv page View PDF GitHub 70 Add to collection Community gpengzhi Paper submitter about 23 hours ago • edited about 23 hours ago Try the live demo: https://huggingface.

co/spaces/xiaomi-research/milmmt-46-translation Reply librarian-bot about 2 hours ago This is an automated message from the Librarian Bot.I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API PAMT: Process-Aligned Reinforcement Learning for Multi-Domain Machine Translation (2026) Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation (2026) Reasoning Before Translation: Enhancing Legal Machine Translation with Structured Reasoning (2026) CAT-Translate: Building Compact Open-Source Models for Japanese-English Translation (2026) MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages (2026) Distill Where the Student Goes: Teacher-Regularized RL for English-Evidence Cross-Lingual RAG (2026) Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR (2026) Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend Reply EditPreview Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images Comment · Sign up or log in to comment Upvote 10 Get this paper in your agent: hf papers read 2608.10812 Don't have the latest CLI?curl -LsSf https://hf.co/cli/install.

sh | bash Models citing this paper 3 Datasets citing this paper 0 No dataset linking this paper Cite arxiv.org/abs/2608.10812 in a dataset README.md to link it from this page.Spaces citing this paper 3 Collections including this paper 2

Related

相關文章

IT之家模型更新

消息稱字節整合 AI 生產力:TRAE、釦子併入豆包,將推統一辦公品牌“豆包工作”

作者:沁滄(實習) 責編:沁滄 評論: 感謝網友 HH_KK 的線索投遞!8 月 24 日消息,據智能湧現消息,字節跳動對旗下的辦公 AI 產品完成了一輪團隊整合:TRAE、釦子(Coze)團隊將整體併入豆包體系,其中 TRAE Work、釦子將與豆包在工作場景的產品能力進行整合;TRAE IDE 及 CLI 將作為豆包品牌下的編程產品線持續發展。

剛剛
鈦媒體模型更新

DeepSeek Harness來了:AI開始製造AI了?

DeepSeek Harness 正式推出,這項新工具被視為 AI 發展的重要里程碑,可能讓 AI 系統具備自主開發或優化其他 AI 的能力。外界關注此技術是否象徵 AI 開始「製造」AI,並可能加速人工智慧的進化與應用。目前相關細節與實際影響仍待進一步觀察。

剛剛
鈦媒體模型更新

AI辦公助手,沒有葵花寶典:五款應用萬字實測報告

AGI-Signal2026.08.24 09:12 · 來自北京全文11936字單項冠軍各有其人。2026年上半年,AI辦公賽道發生了一個根本性變化,工具不再滿足於當“對話框”,而是試圖接管完整任務,寫一段文案、做完一份報告、生成一份PPT,甚至跨應用操作。

37 分鐘前

字節整合AI辦公產品,TRAE、釦子團隊併入豆包

其中,TRAE Work、釦子將與豆包的工作場景產品能力整合;TRAE IDE及CLI則作為豆包品牌下的編程產品線繼續發展。調整後,相關產品和運營團隊統一向豆包產品負責人趙祺彙報。(iFeng Tech)TRAE與釦子此前均隸屬於字節跳動產品研發和工程架構部,前者最初定位AI編程產品,後者則聚焦AI智能體開發平臺,並持續探索不同Agent方向。

44 分鐘前6100