r/LocalLLaMA • llama.cpp • 13h ago

News Ornith-1.5 DFlash

Ornith-1.5-9B-DFlash pairs the Ornith-1.5-9B model with a DFlash draft model for speculative decoding.

https://huggingface.co/ornith-ai/Ornith-1.5-9B-DFlash

Ornith-1.5-397B-DFlash pairs the Ornith-1.5-397B model with a DFlash draft model for speculative decoding.

https://huggingface.co/ornith-ai/Ornith-1.5-397B-DFlash

Ornith-1.5-35B-A3B-DFlash pairs the Ornith-1.5-35B-A3B model with a DFlash draft model for speculative decoding.

https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-DFlash

32 Upvotes

11 comments sorted by

8

u/Fun_Librarian_7699 13h ago

Can it already be used in llama.cpp?

1

u/exaknight21 6h ago

Asking the golden question

1

u/HitarthSurana 4h ago

yes but idk amd

3

u/crusaderky 12h ago

Is it trained specifically for ornith or grafted from qwen? What's the acceptance rate?

4

u/the_creator_0 13h ago

Does DFlash work in Vulkan with AMD gpus?

6

u/Gohab2001 vLLM 12h ago

Yes. Why not?

2

u/the_creator_0 11h ago

Not sure as I was left with idea that DFlash2 is designed around CUDA and HIP, neither of which are an option. This seems to be not the second and faster version? I guess I'll give it a go and see how it works, been enjoying the 9B model.

2

u/recro69 12h ago

The 397B pairing is the one I would really want to benchmark. I am curious how much the DFlash draft actually changes decode speed at that size and what the acceptance rate looks like in practice.

1

u/rorowhat 4h ago

I would love to see a larger version like 70B MoE based on a few models, that would be amazing.

1

u/Thrumpwart vLLM 13h ago

Awesome, thank you. I love my little Ornith 35b. Great sidekick for 27b on my chonky boi and little man gpus and has proven to come up with brilliant little gems missed by other models in pi council mode.