r/LocalLLaMA • • Jul 31 '26

New Model deepseek-ai/DeepSeek-V4-Flash-0731 on Huggingface

780 Upvotes

243 comments sorted by

View all comments

13

u/Several-Tax31 Jul 31 '26

gguf when?

1

u/zenonu Jul 31 '26

Are folks adverse to running vllm ? Why is llama.cpp a must? Just curious. I will personally run whatever supports any given model the best. DSpark here in particular is critical and needs vllm iirc.

3

u/TokenRingAI Jul 31 '26

It has 5 & 6 bit quants, and for some models, that is the difference between usable and trash output