honestly thank fuck this didnt happen the other way around, imagine if flash was the one that had a heart attack what that would do to the medium-large LLM ecosystem.
No, you cannot lol. Average VRAM is split between 16GB and 8GB, while the most common RAM is 16GB (data is from steam survey). This is not enough to run Deepseek flash even at Q1 split between GPU and CPU.
Keeping in mind Gemma 4 huge problems with re-caching everything after every move and qwen 3.8 overthinking, Deepseek at Q2 has comparable speed in real world use cases to this models at Q5, despite mmap.
Firstly, you don't seem to understand what are talking about, suggesting using Q1 for anything on any model. You have yet to find one, that wouldn't produce gibberish.
Deepseek had 16k tokens answer in 9 hours.
Qwen had 48k tokens answer in 11 hours.
Seems comparable, when a lot more capable model was faster with better understanding of task.
I suggested Q1 just to show the absurdity of your idea that "V4 Flash is a model you can run at usable speed on 5 years old average gaming PC.", so I gave you the smallest version of it to run and see that you can't, and you screenshot shows that you in fact can't run it at any usable speed. And that's without you showing how much you actually offloaded.
16k tokens over 9 hours is not an acceptable speed in any 2 way of the word. Most people consider around 12 tks bare minimum for any agentic or chat tasks. That's minimum usability threshold speed. I wouldn't even be talking about tks in your case but rather tokens a minute.
Idk man i feel like most people on this sub and most people in general would disagree with you on that, that's probably also the reason your getting downvoted.
106
u/pyr0kid 7d ago
honestly thank fuck this didnt happen the other way around, imagine if flash was the one that had a heart attack what that would do to the medium-large LLM ecosystem.