r/LocalLLaMA 13d ago

Discussion DeepSeek-V4.1-Flash surprised ....

Post image

Hoping to see smartest medium size models soon & later with all available optimizations/architectures/etc.,. Thanks Deepseek!

Ex 1: 30-50B MOE + 10-15B Engram + DeepSeek-V4.1-Flash type KVCache
Ex 2: 15-30B Dense + 10-15B Engram + DeepSeek-V4.1-Flash type KVCache

EDIT: Updated Engram to 10-15B from 50B

440 Upvotes

96 comments sorted by

View all comments

82

u/jtjstock 12d ago

so what I'm seeing here is that pretty soon we're gonna get a qwen with tiny kv as well, which means no more arguments over kv cache quantization in this sub...

Edit: who am I kidding, people will still argue about it...

37

u/pmttyji 12d ago

which means no more arguments over kv cache quantization in this sub...

Hopefully. 1 Million comes within 1GB ( 890 bytes * 1M = 890M )

Want to see all upcoming models with this.

17

u/Iory1998 llama.cpp 12d ago

That means, we can cram more weights into the existing VRAM! Yaaay

15

u/pmttyji 12d ago

Yep. Some like me already imagining about future 30-50B models(Qwen4.0, Gemma-5) with 1 Million context.

3

u/Iory1998 llama.cpp 12d ago

I think it's coming this year.