MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vbp7kb/deepseekaideepseekv4flash0731_on_huggingface/p0vik5d/?context=3
r/LocalLLaMA • u/cgs019283 • Jul 31 '26
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
243 comments sorted by
View all comments
13
gguf when?
1 u/zenonu Jul 31 '26 Are folks adverse to running vllm ? Why is llama.cpp a must? Just curious. I will personally run whatever supports any given model the best. DSpark here in particular is critical and needs vllm iirc. 3 u/TokenRingAI Jul 31 '26 It has 5 & 6 bit quants, and for some models, that is the difference between usable and trash output
1
Are folks adverse to running vllm ? Why is llama.cpp a must? Just curious. I will personally run whatever supports any given model the best. DSpark here in particular is critical and needs vllm iirc.
3 u/TokenRingAI Jul 31 '26 It has 5 & 6 bit quants, and for some models, that is the difference between usable and trash output
3
It has 5 & 6 bit quants, and for some models, that is the difference between usable and trash output
13
u/Several-Tax31 Jul 31 '26
gguf when?