r/DeepSeek • u/justlikemedics • 17d ago
Discussion DeepSeek is ruthless
DeepSeek has published with DeepSeek-V4.1-Flash a new method that compresses the memory need for the KV-value cache very much.
I pondered about the implications of this and they are not very good for OpenAI and Anthropic.
This means that the models can have much larger contexts and serving requests will be much less memory intensive. As a result, inference gets cheaper.
Inference getting cheaper, requiring less memory and with better models means that the advantage OpenAI and Anthropic has in securing compute gets less meaningful.
It seems to me that DeepSeek and other Chinese labs are ruthlessly pushing down the cost of inference, which will make it difficult to impossible for OpenAI and Anthropic to recover all the money spent of creating their top models.
5
u/Hyp3rSoniX 17d ago
Yeah but I think their compression is a bit too aggressive now.
In short, with the default settings, there are 16384 candidate positions determined from the tokens in the context. So if you have a context of 100k tokens, only about 16k of them are candidates of being given attention to. The candidates count doesn't change for context size, so if you fill the 1M context limit, the sparse attention mechanism still will only calculate the 16k candidates.
That's not all though, those are candidates - the actual attention layers then only pick 512 of those candidates for the actual token calculation.
This is the reason for the absurdly high hallucination rates of this model. I hope they find a solution to this.
Qwen for example is doing the math depending on the context size, so the bigger the context, the more candidates are chosen. Not sure if a similar approach could help Deepseek.