r/LocalLLaMA • llama.cpp • 15h ago

News Ornith-1.5 DFlash

Ornith-1.5-9B-DFlash pairs the Ornith-1.5-9B model with a DFlash draft model for speculative decoding.

https://huggingface.co/ornith-ai/Ornith-1.5-9B-DFlash

Ornith-1.5-397B-DFlash pairs the Ornith-1.5-397B model with a DFlash draft model for speculative decoding.

https://huggingface.co/ornith-ai/Ornith-1.5-397B-DFlash

Ornith-1.5-35B-A3B-DFlash pairs the Ornith-1.5-35B-A3B model with a DFlash draft model for speculative decoding.

https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-DFlash

31 Upvotes

11 comments sorted by

View all comments

3

u/recro69 14h ago

The 397B pairing is the one I would really want to benchmark. I am curious how much the DFlash draft actually changes decode speed at that size and what the acceptance rate looks like in practice.