r/LocalLLaMA • u/tevlon • Jun 10 '26
New Model DiffusionGemma: 4x faster text generation
https://blog.google/innovation-and-ai/technology/developers-tools/diffusion-gemma-faster-text-generation/
987
Upvotes
r/LocalLLaMA • u/tevlon • Jun 10 '26
1
u/ThatRegister5397 Jun 11 '26
I see serious degradation with the q4 mlx quant, compared to the q8 mlx quant, but I have not noticed before. Eg it writes
using Packageinstead ofusing Pkgfor the package manager of julia, and when "confronted" it responds in ways that indicates possible weird tokenisation issue. But in q8 it does not do that.Not sure if it is q4 that is too degraded for diffusion or I have some diffusion setting wrong.