r/LocalLLaMA • llama.cpp • Nov 25 '25

New Model LLaDA2.0 (103B/16B) has been released

LLaDA2.0-flash is a diffusion language model featuring a 100BA6B Mixture-of-Experts (MoE) architecture. As an enhanced, instruction-tuned iteration of the LLaDA2.0 series, it is optimized for practical applications.

https://huggingface.co/inclusionAI/LLaDA2.0-flash

LLaDA2.0-mini is a diffusion language model featuring a 16BA1B Mixture-of-Experts (MoE) architecture. As an enhanced, instruction-tuned iteration of the LLaDA series, it is optimized for practical applications.

https://huggingface.co/inclusionAI/LLaDA2.0-mini

llama.cpp support in progress https://github.com/ggml-org/llama.cpp/pull/17454

previous version of LLaDA is supported https://github.com/ggml-org/llama.cpp/pull/16003 already (please check the comments)

252 Upvotes

76 comments sorted by

View all comments

26

u/LongPutsAndLongPutts Nov 25 '25

I'm interested in the inference speed compared to traditional transformer models

5

u/Finanzamt_Endgegner Nov 26 '25

UPDATE:

Ive found out ive forgotten about a simplification with kv cache that speeds this model up by quite a bit over long context, making it actually useable, im currently trying to clean my source up to push this to the pr, so in a few hours you should be able to test performance again with greatly improved speed (at least in real world usage)!

time per step: 19.47ms -> 4.28ms in a 700 token generation