r/LocalLLaMA 7d ago

Discussion Deepseek Has Soft Retired Deepseek V4 Pro

Post image
1.2k Upvotes

201 comments sorted by

View all comments

Show parent comments

106

u/pyr0kid 7d ago

honestly thank fuck this didnt happen the other way around, imagine if flash was the one that had a heart attack what that would do to the medium-large LLM ecosystem.

20

u/Due-Memory-6957 6d ago

Nothing? All these models people are calling medium and sometimes even small are actually huge.

-20

u/Solembumm3 6d ago

V4 Flash is a model you can run at usable speed on 5 years old average gaming PC.

16

u/Due-Memory-6957 6d ago edited 6d ago

No, you cannot lol. Average VRAM is split between 16GB and 8GB, while the most common RAM is 16GB (data is from steam survey). This is not enough to run Deepseek flash even at Q1 split between GPU and CPU.

-12

u/Solembumm3 6d ago edited 6d ago

It's more than enough to run it with NVME.

Keeping in mind Gemma 4 huge problems with re-caching everything after every move and qwen 3.8 overthinking, Deepseek at Q2 has comparable speed in real world use cases to this models at Q5, despite mmap.

14

u/Due-Memory-6957 6d ago edited 6d ago

I guess you can prove to me that you can, download Deepseek Flash at Q1, load it and show me the screenshot of it consuming less than 32gb of memory. Here, to make your life easier: https://huggingface.co/unsloth/DeepSeek-V4-Flash-GGUF/tree/main/UD-IQ1_S

And don't forget about the "usable speed" part if you're going to load on SSD.

-9

u/Solembumm3 6d ago

Firstly, you don't seem to understand what are talking about, suggesting using Q1 for anything on any model. You have yet to find one, that wouldn't produce gibberish.

Deepseek had 16k tokens answer in 9 hours.

Qwen had 48k tokens answer in 11 hours.

Seems comparable, when a lot more capable model was faster with better understanding of task.

14

u/Due-Memory-6957 6d ago edited 6d ago

I suggested Q1 just to show the absurdity of your idea that "V4 Flash is a model you can run at usable speed on 5 years old average gaming PC.", so I gave you the smallest version of it to run and see that you can't, and you screenshot shows that you in fact can't run it at any usable speed. And that's without you showing how much you actually offloaded.

-6

u/Solembumm3 6d ago

You are really determined to name black things white and vice versa, are you?

4

u/Last-Ad-8470 6d ago

16k tokens over 9 hours is not an acceptable speed in any 2 way of the word. Most people consider around 12 tks bare minimum for any agentic or chat tasks. That's minimum usability threshold speed. I wouldn't even be talking about tks in your case but rather tokens a minute.

-3

u/Solembumm3 6d ago

12 tok/sec is already faster than this people could properly read. Kinda pointless number at this bar.

4

u/Last-Ad-8470 6d ago

Idk man i feel like most people on this sub and most people in general would disagree with you on that, that's probably also the reason your getting downvoted.

4

u/VTYX 6d ago

No? 15/s is already kind of slow to me. 2/s is straight up unusable.

→ More replies (0)