谷歌改造文本擴散模型
Ad Skip to content Google's DiffusionGemma proves you don't need to train from scratch to build a text diffusion model Jonathan Kemper View the LinkedIn Profile of Jonathan Kemper Aug 9, 2026 Nano Banana Pro prompted by THE DECODER Instead of training a new model from scratch, Google DeepMind retrofitted Gemma 4 into a diffusion model.
The newly published report explains how it works and where the tradeoffs are.Google DeepMind released DiffusionGemma as a model in mid-June and has now followed up with the technical report.
Unlike standard language models that generate text one token at a time, DiffusionGemma refines blocks of 256 tokens in parallel, similar to how image AIs pull a picture out of noise.On an Nvidia H100 accelerator, the model hits about 1,500 tokens per second.
Building a new model from scratch wasn't necessary.The team started with the existing Gemma-4-26B-A4B and converted it into a diffusion model using less than ten percent of the original training token budget, according to the report.
DiffusionGemma delivers several times the output speed of the Gemma 4 models and previous diffusion models while maintaining comparable accuracy.
| Image: Google Two training stages balance quality and speed In the first of two steps, the model learns to reconstruct noisy text blocks from example data.A combined phase of reinforcement learning and sampler distillation follows, which Google calls SD·RL.
Reinforcement learning typically boosts answer quality, while sampler distillation lets the model get by with fewer compute steps.Google merges both into a single process.Google DeepMind doesn't train DiffusionGemma from scratch but converts the finished Gemma 4 model through two training stages.
| Image: Google According to the report, this combined approach raises quality on reasoning benchmarks by an average of ten points while nearly quadrupling the number of tokens per compute step.As a side effect, DiffusionGemma's answers run about 50 percent shorter, which further boosts speed.
Bidirectional reasoning lets the model correct itself Standard language models have to commit to the first digit of an answer before they've worked through the reasoning.
In a math problem from the report, Gemma 4 starts its response with "-1," realizes during its derivation that "-25" is correct, and tacks on a correction afterward.DiffusionGemma develops the answer and reasoning in parallel, so it can fix mistakes before the output is finalized.
Unlike an autoregressive model, DiffusionGemma can correct an early wrong answer during later denoising steps.| Image: Google Sudoku solving works on the same principle, since every entry depends on entries that come later.
After minimal fine-tuning, DiffusionGemma solves close to 85 percent of puzzles correctly, while the base model fails at the task entirely.
Structured outputs like JSON or code repairs finish after just two to three refinement steps, according to the report, because the input already determines most tokens.
DiffusionGemma also keeps its original ability to generate text word by word, letting users switch between both modes depending on the task.DiffusionGemma trails the autoregressive Gemma 4 on quality benchmarks but leads on output speed at about 1,500 tokens per second.
| Image: Google Reasoning gaps and multi-user limits persist Absolute performance falls short of the autoregressive base model.Google points to several reasons for this.DiffusionGemma wasn't trained as a diffusion model from the start but was retrofitted after the fact.
The subsequent training phase was relatively short, and the second step, SD·RL, prioritized speed over peak quality.The architecture, training data, and other settings were also carried over from the original Gemma 4 model, which aren't necessarily ideal for diffusion.
The model occasionally gets stuck in repetition loops, producing individual words multiple times in a row.This is an artifact of the aggressively reduced compute steps.
On multimodal tasks, DiffusionGemma sometimes forgets to close its reasoning section properly, which artificially drags down benchmark scores.The speed advantage also holds mainly for single-user scenarios.
Once about 32 concurrent requests hit the model, standard language models catch up on throughput.
Google explicitly calls DiffusionGemma an experimental model and says the release is meant to speed up research on text diffusion while giving the community a foundation for specialized, resource-efficient adaptations.
The model is already being used by the startup Interfaze for multilingual speech recognition and in a research project on interactive radiology report generation.Google previously made the model available under an Apache 2.0 license on Hugging Face.
Its predecessor is Gemini Diffusion, which Google demoed in May 2025.AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Subscribe now Read on for the full picture.Subscribe for hype-free coverage.Access to all THE DECODER articles.Read without distractions – no Google ads.Access to comments and community discussions.Weekly AI newsletter.6 times a year: “AI Radar” – deep dives on key AI topics.
Up to 25 % off on KI Pro online events.Access to our full ten-year archive.Get the latest AI news from The Decoder.Subscribe to The Decoder BETA-TEST × wpDiscuzInsert BETA-TEST × wpDiscuzInsert
Related
相關文章

AI偶像,不能照搬真人明星的邏輯
虛眸2026.08.24 14:26 · 來自北京全文4075字00:00 / 11:50就算看著再逼真,也知道不是人。文 | 虛眸第一波AI明星出道,並不順利。AI短劇《被裁掉的女孩》的虛擬女主角方桃子,代言隱形眼鏡時稱“戴了一天很舒服”,隨即被大眾質疑:一個沒有身體的AI角色,如何感受“舒服”?《與你深情,侵入餘生》中的男女主段宴和容寄僑,以演員身份二搭“出演”新劇《分手後男頻女頻大亂鬥》,讓粉絲對自家偶像究竟是誰、屬於哪個次元的世界產生認知混亂。

小米米家智能魚缸 2 Pro 開啟眾籌:支持自動餵食,眾籌價 599 元
米家智能魚缸 2 Pro 今日在小米有品開啟眾籌,售價 599 元。產品配備定製循環水路、自動餵食器及 1.47 寸 LCD 彩屏,支持米家 App 遠程操控與小愛同學語音控制,還可根據魚種信息智能匹配運行策略。#小米有品# 你心動了嗎?

小米推出米家掃拖機器人 7C:滾筒活水增壓拖地,到手價 2069.1 元
作者:浩渺 責編:浩渺 評論: 感謝網友 很宅很怕生 的線索投遞!8 月 23 日消息,米家掃拖機器人 7C 現已在小米有品上架預約,官方標價 2299 元,券後到手價 2069.1 元,8 月 26 日 10 點現貨開售(點擊前往)。從商品頁面獲悉,這款新品支持滾筒活水拖地,配備恆壓恆溼滾筒拖布,200 轉 / 分鐘高速洗拖,強效清潔咖啡漬、油漬、寵物腳印等頑固汙漬; 配合活水邊拖邊洗,有效減少二次汙染。
Building an End-to-End Document Intelligence Pipeline with deepDoctection
In this tutorial, we implement a document intelligence pipeline with deepDoctection 1.2.x that combines layout detection, table structure recognition, OCR, reading-order reconstruction, annotation linking, and structured export in a single workflow.

抖快B紅集體押注“AI互動內容”,創作者如何抓住新機會?
中國四大內容平台抖音、快手、B站與小紅書近期同時押注AI互動內容,讓觀眾能與AI角色對話或主導劇情,被視為下一波短影音戰場的提前卡位。海外相關新創App也獲得近億美元融資,使這場競賽升級為資金與流量的軍備賽。對創作者而言,這波浪潮將編劇與產品設計能力納入創作門檻,平台雖推出工具與分成機制,但商業模式仍在摸索階段。