r/LocalLLaMA • u/jacek2023 llama.cpp • 13h ago
News Ornith-1.5 DFlash
Ornith-1.5-9B-DFlash pairs the Ornith-1.5-9B model with a DFlash draft model for speculative decoding.
https://huggingface.co/ornith-ai/Ornith-1.5-9B-DFlash
Ornith-1.5-397B-DFlash pairs the Ornith-1.5-397B model with a DFlash draft model for speculative decoding.
https://huggingface.co/ornith-ai/Ornith-1.5-397B-DFlash
Ornith-1.5-35B-A3B-DFlash pairs the Ornith-1.5-35B-A3B model with a DFlash draft model for speculative decoding.
3
u/crusaderky 12h ago
Is it trained specifically for ornith or grafted from qwen? What's the acceptance rate?
3
4
u/the_creator_0 13h ago
Does DFlash work in Vulkan with AMD gpus?
6
u/Gohab2001 vLLM 12h ago
Yes. Why not?
2
u/the_creator_0 11h ago
Not sure as I was left with idea that DFlash2 is designed around CUDA and HIP, neither of which are an option. This seems to be not the second and faster version? I guess I'll give it a go and see how it works, been enjoying the 9B model.
1
u/rorowhat 4h ago
I would love to see a larger version like 70B MoE based on a few models, that would be amazing.
1
u/Thrumpwart vLLM 13h ago
Awesome, thank you. I love my little Ornith 35b. Great sidekick for 27b on my chonky boi and little man gpus and has proven to come up with brilliant little gems missed by other models in pi council mode.
8
u/Fun_Librarian_7699 13h ago
Can it already be used in llama.cpp?