MiLMMT翻譯提質

2026年8月13日 00:00
站內 AI 整理稿

Papers arxiv:2608.

10812 Copy markdown Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation Published on Aug 11 · Submitted by Pengzhi Gao on Aug 12 · Xiaomi Research Upvote 10 +2 Authors: Chris Han ,Pengzhi Gao ,Pei Fu ,Jian Luan Abstract Open multilingual translation models are improved via group relative policy optimization with reference-free quality rewards and checkpoint interpolation, surpassing strong open and proprietary baselines.

Generated by thinkingmachines/Inkling-Small We study reference-free post-training for multilingual machine translation with open large language models.Starting from the supervised-finetuned MiLMMT-46-v0.

1 models, we apply Group Relative Policy Optimization (GRPO) with a reward that averages two reference-free quality estimation models and is gated by language identification.

We then linearly interpolate the supervised fine-tuning (SFT) and reinforcement learning (RL) model checkpoints to obtain MiLMMT-46-v1.0.

Across 46 languages, the resulting models consistently improve translation quality over their SFT counterparts, outperform strong recent open baselines, including Seed-X, HY-MT2, and TranslateGemma, and achieve leading reference-free scores against evaluated proprietary systems such as Google Translate, Gemini 3 Pro, and GPT-5.

We further investigate on-policy distillation and find that it reaches, but does not surpass, the quality frontier achieved by RL with checkpoint interpolation.We release the models and code to facilitate future research.

View arXiv page View PDF GitHub 70 Add to collection Community gpengzhi Paper submitter about 23 hours ago • edited about 23 hours ago Try the live demo: https://huggingface.

co/spaces/xiaomi-research/milmmt-46-translation Reply librarian-bot about 2 hours ago This is an automated message from the Librarian Bot.I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API PAMT: Process-Aligned Reinforcement Learning for Multi-Domain Machine Translation (2026) Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation (2026) Reasoning Before Translation: Enhancing Legal Machine Translation with Structured Reasoning (2026) CAT-Translate: Building Compact Open-Source Models for Japanese-English Translation (2026) MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages (2026) Distill Where the Student Goes: Teacher-Regularized RL for English-Evidence Cross-Lingual RAG (2026) Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR (2026) Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend Reply EditPreview Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images Comment · Sign up or log in to comment Upvote 10 Get this paper in your agent: hf papers read 2608.10812 Don't have the latest CLI?curl -LsSf https://hf.co/cli/install.

sh | bash Models citing this paper 3 Datasets citing this paper 0 No dataset linking this paper Cite arxiv.org/abs/2608.10812 in a dataset README.md to link it from this page.Spaces citing this paper 3 Collections including this paper 2

Related

相關文章

鈦媒體模型更新

AI辦公助手,沒有葵花寶典:五款應用萬字實測報告

AGI-Signal2026.08.24 09:12 · 來自北京全文11936字單項冠軍各有其人。2026年上半年,AI辦公賽道發生了一個根本性變化,工具不再滿足於當“對話框”,而是試圖接管完整任務,寫一段文案、做完一份報告、生成一份PPT,甚至跨應用操作。

剛剛

Anthropic新模型偷「吃瓜」,最強Fable 5爆冷

Anthropic 近日推出新款 AI 模型,在內部測試中意外展現「吃瓜」能力,引發社群熱議。該模型不僅能快速理解網路迷因與流行語,更在特定任務上表現出人意料,讓原本被外界視為最強對手的 Fable 5 爆冷落後,業界對這項結果感到相當驚訝。目前 Anthropic 官方尚未針對模型實際表現與測試細節做出完整說明,市場則持續關注後續可能的技術更新與應用方向。

剛剛

Kimi K2.5 月底退役:月之暗面第一代萬億參數多模態模型謝幕

月之暗面官宣第一代萬億參數多模態模型Kimi K2.5將於本月底結束服役。該模型今年1月推出並開源,是Kimi迄今最全能模型,採用原生多模態架構,支持視覺與文本輸入、思考/非思考模式、對話與Agent任務,在Agent、代碼、圖像、視頻及通用智能取得開源SOTA。K3將接力,參數規模再上臺階。

14 分鐘前4700
雷峰網模型更新

光子躍遷亮相BIRTV 2026:以"AI+影像"重構創作範式,三大板塊解碼下一代影像生態

8月19日,BIRTV 2026(北京國際廣播電影電視展覽會)在北京拉開帷幕。在這場匯聚全球廣電與影像領域頂尖技術與創意的盛會上,光子躍遷以"AI+影像"為核心敘事,攜個人智能影像生態重磅亮相,向行業展示了一個由AI驅動、以人為中心的影像未來。與行業展會常見的深色科技風不同,光子躍遷的展臺以純淨白色為主基調,輔以品牌藍色進行點睛點綴,在千篇一律的深色展臺中脫穎而出,傳遞出品牌年輕、活力、面向未來的基因。

5 小時前