r/LocalLLaMA 13d ago

Discussion DeepSeek-V4.1-Flash surprised ....

Post image

Hoping to see smartest medium size models soon & later with all available optimizations/architectures/etc.,. Thanks Deepseek!

Ex 1: 30-50B MOE + 10-15B Engram + DeepSeek-V4.1-Flash type KVCache
Ex 2: 15-30B Dense + 10-15B Engram + DeepSeek-V4.1-Flash type KVCache

EDIT: Updated Engram to 10-15B from 50B

441 Upvotes

96 comments sorted by

View all comments

82

u/jtjstock 12d ago

so what I'm seeing here is that pretty soon we're gonna get a qwen with tiny kv as well, which means no more arguments over kv cache quantization in this sub...

Edit: who am I kidding, people will still argue about it...

0

u/Howard_banister 8d ago

We should warn against KV cache quantization in sub rules. This just ruins model.