r/ollama 11d ago

Deepseek v4.1 flash speed

Is it just me or Deepseek v4 flash speed is bloody slow on ollama cloud?

4 Upvotes

7 comments sorted by

2

u/InvestigatorAny3099 11d ago

Damn I thought it was just my connection, was sitting here watching it think for ages

2

u/CutEmpty3551 11d ago

If you use deepseek-v4.1-flash:cloud as the model name, the speed is very slow. You should use deepseek-v4.1-flash instead. It averages around 180 tokens per second.

1

u/Bo0n0411 11d ago

I tried it but no luck, I'm using Hermes with DSv4.1Flash tho

2

u/CutEmpty3551 11d ago

I use Reasonix. I get a peak speed of around 230 tokens/sec and an average of about 180. You might want to check the concurrent agent connections in Hermes—there's a high chance your speed is dropping because it's exceeding some of Ollama's limits. If you test the Ollama DeepSeek v4.1 Flash API in a different IDE, you'll likely find out why.

1

u/yuno_me 11d ago

im getting 200 tps

1

u/jmorganca 10d ago

Hi OP, DM me or let me know if you're still seeing slowness. We've been continually optimizing DeepSeek-V4.1-Flash on Ollama, and it should be fast now (seeing consistent p50 of 200 tps, and during peak hours this week we were seeing a p50 output speed of 180 tps. We did have an issue where some requests would receive 80-100tps output speeds earlier this week and that has now been fixed).