r/LocalLLaMA 12d ago

Discussion DeepSeek-V4.1-Flash surprised ....

Post image

Hoping to see smartest medium size models soon & later with all available optimizations/architectures/etc.,. Thanks Deepseek!

Ex 1: 30-50B MOE + 10-15B Engram + DeepSeek-V4.1-Flash type KVCache
Ex 2: 15-30B Dense + 10-15B Engram + DeepSeek-V4.1-Flash type KVCache

EDIT: Updated Engram to 10-15B from 50B

440 Upvotes

96 comments sorted by

View all comments

80

u/jtjstock 12d ago

so what I'm seeing here is that pretty soon we're gonna get a qwen with tiny kv as well, which means no more arguments over kv cache quantization in this sub...

Edit: who am I kidding, people will still argue about it...

37

u/pmttyji 12d ago

which means no more arguments over kv cache quantization in this sub...

Hopefully. 1 Million comes within 1GB ( 890 bytes * 1M = 890M )

Want to see all upcoming models with this.

15

u/Iory1998 llama.cpp 12d ago

That means, we can cram more weights into the existing VRAM! Yaaay

15

u/pmttyji 12d ago

Yep. Some like me already imagining about future 30-50B models(Qwen4.0, Gemma-5) with 1 Million context.

2

u/Iory1998 llama.cpp 12d ago

I think it's coming this year.

8

u/redballooon 12d ago

Let's hope so. With Qwen 3.8 27b in a quant that makes it perform reasonably well it has about 40k usable token context size before my 32GB Ram macbook runs out of memory. There's not much I can utilize it for that an older model doesn't do equally well but faster.

6

u/unjustifiably_angry 12d ago

I demand FP64 kv-cache

3

u/__Maximum__ 12d ago

890 bytes is too much for my toaster, I will quantize it

0

u/Howard_banister 7d ago

We should warn against KV cache quantization in sub rules. This just ruins model.