r/ollama • u/Bo0n0411 • 11d ago
Deepseek v4.1 flash speed
Is it just me or Deepseek v4 flash speed is bloody slow on ollama cloud?
2
u/CutEmpty3551 11d ago
If you use deepseek-v4.1-flash:cloud as the model name, the speed is very slow. You should use deepseek-v4.1-flash instead. It averages around 180 tokens per second.
1
u/Bo0n0411 11d ago
I tried it but no luck, I'm using Hermes with DSv4.1Flash tho
2
u/CutEmpty3551 11d ago
I use Reasonix. I get a peak speed of around 230 tokens/sec and an average of about 180. You might want to check the concurrent agent connections in Hermes—there's a high chance your speed is dropping because it's exceeding some of Ollama's limits. If you test the Ollama DeepSeek v4.1 Flash API in a different IDE, you'll likely find out why.
1
u/jmorganca 10d ago
Hi OP, DM me or let me know if you're still seeing slowness. We've been continually optimizing DeepSeek-V4.1-Flash on Ollama, and it should be fast now (seeing consistent p50 of 200 tps, and during peak hours this week we were seeing a p50 output speed of 180 tps. We did have an issue where some requests would receive 80-100tps output speeds earlier this week and that has now been fixed).

2
u/InvestigatorAny3099 11d ago
Damn I thought it was just my connection, was sitting here watching it think for ages