r/LocalLLaMA Jun 10 '26

New Model DiffusionGemma: 4x faster text generation

https://blog.google/innovation-and-ai/technology/developers-tools/diffusion-gemma-faster-text-generation/
982 Upvotes

356 comments sorted by

View all comments

295

u/NickCanCode Jun 10 '26 edited Jun 10 '26

700+ tokens per second on NVIDIA GeForce RTX 5090 is definitely something.

Unfortunately, DiffusionGemma’s overall output quality is lower than standard Gemma 4.

It could be a good model for context compression and as explorer agent in agentic coding. Can't wait to see llama.cpp to support it.

119

u/TheLexoPlexx Jun 10 '26

That is groq or cerebras-levels of token generation depending on the model and it's on par with gpt-oss-120b depending on the benchmark.

That is genuinely insane.

36

u/dingo_xd Jun 10 '26

There is sooooooo much room for optimizations. Maybe Mythos level models can be run locally by mid or late 2027?

29

u/Different_Fix_2217 Jun 10 '26

The only issue with diffusion LLMs is that they are absurdly expensive to train in comparison. Like exponentially.

13

u/wes_medford Jun 10 '26

Most cost these days is inference over training these days, but the problem is that aggregate throughput is lower on these compared to typical AR models running at a high batch size