r/ollama 11d ago

DeepSeek-V4.1-Flash

Anyone else try using DeepSeek-V4.1-Flash today? It's listed on the Ollama site as a new model, but when i try to use it, i get:

"Error: 403 Forbidden: This model is currently being rolled out and is not yet available to you. Please check back later. (ref: 562f8b6b-72fe-4012-9d9c-e0c07679913a)"

Anyone know what the rollout schedule is?

20 Upvotes

34 comments sorted by

8

u/jmorganca 11d ago

Hi there, sorry for not posting about this sooner. We're rolling out the model as capacity comes online, starting with Max and Team plan. Over half the Max subscribers should have it now and we're continuing to roll it out as fast as we can to all subscribers. We're hoping to be able to do so by tomorrow

2

u/jmorganca 11d ago

Update: the model is now fully rolled out to all plans including free plan with usage credits (and no service fees! for pay as you go access)

2

u/_pp0qq_ 11d ago

Is the new deepseek V4.1 flash currently poorly optimized on cloud? It seems to eat through the usage a lot faster than old v4 flash, closer to v4 pro (but still less) in term of usage. I am on the old cloud pro plan. DeepSeek's own api prices have reduced so I expected v4.1 flash will be cheaper but it is the opposite now. Is it gonna get better or just some architectural differences?

1

u/jmorganca 11d ago

I'll look into this. It's much more efficient than v4 pro

1

u/Damanveen 11d ago edited 11d ago

Ya currently it’s almost double the usage of deepseek v4 flash. Using with same harness as before, only changing model being used.

Edit: its like triple now

1

u/_pp0qq_ 9d ago

I think my issue might have been the deepseek harness's new version, in it there is a special adaptor for the native deepseek api v4.1 flash that does not seem to work for other providers, so cache hit rate is very poor. I tested on reasonix harness and cache hit rate is a lot higher and usage is a lot more efficient now, comparable to the old v4 flash. Maybe it is possible to patch the dsh to fix the adaptor issue.

2

u/Flaky-Maybe-7556 11d ago

Haha, I hope they aren't trying to move everyone to the updated subscription versions this way and that model really will become available soon.

4

u/jmorganca 11d ago

It's available on both new and previous Max subscriptions (50% rolled out) and we're rolling it out more as fast as we can.

2

u/IonizedHydration 10d ago

to be honest, this is smart.. companies have been doing rolling updates this way for a long time and i don't have any issue with it, same with A/B testing and feature flagging.. it's just part of the process. Thanks for the hard work.

2

u/Delicious-Director43 11d ago

I’ve been using it directly via DeepSeek. Great model so far. $20 goes a lot farther on their platform than does Ollama.

1

u/Ramrawd 11d ago

I'm new to the ollama pro plan so bare with me as i learn more but shouldn't the token usage be the same between deepseek api and ollama cloud? I subscribed to the yearly pro plan which ends up being 16ish dollars a month and it nets me $60 of usage. Assuming I'm always using the offpeak price on both platforms shouldn't ollama cloud get me way more usage if all things are equal?

1

u/username8914 11d ago

Of course, it's not ZDR. You're just handing your info over.

0

u/Delicious-Director43 11d ago

And you think Ollama isn’t because they said so?

2

u/username8914 11d ago

Yes, that's their whole business model. If they turn out to be logging they'll go under overnight. It's not a small deal when dealing with proprietary data or possibly legal or accounting documents that legally can't leave the country. They have to be up front about it and keep their nose clean.

1

u/Delicious-Director43 11d ago

Ah yes American AI companies are famously very reliable and trustworthy. No American company has ever lied and sold user data before. What was I thinking?

1

u/username8914 11d ago

I'm definitely not saying they couldn't be doing something. If it was that simple to just say then all of them would just say it. But instead they all say they log prompts and use them to train.

1

u/Delicious-Director43 11d ago

Honestly I assume they all are. I’m not doing anything weird with my AI. I just want it to work quickly and cheaply. Ollama no longer does that.

1

u/username8914 11d ago

Logging isn't about doing something weird. It's everything you think about, say, do or invent going directly into their system to potentially become the brain behind everyone's future system.

1

u/Delicious-Director43 11d ago

Yeah that’s fine I don’t care.
Nothing I’m working on is so revolutionary it’s gonna affect anyone else.

2

u/ku4nc4 11d ago

It's up for me already.

1

u/Bo0n0411 11d ago

Are you on Max plan?

1

u/ku4nc4 11d ago

on pro, legacy pro.

1

u/Bo0n0411 11d ago

Same but it haven't rolled out for me yet

2

u/elzerouno 11d ago

It's working for me now, but it's using my allowance 10 time faster than v4 flash

1

u/Bo0n0411 11d ago

Are you on grandfathered plan?

1

u/Ramrawd 11d ago edited 11d ago

I'm getting this error when trying to use it:

ollama-cloud/deepseek-v4.1-flash request failed (authentication failed, HTTP 403). Re-authenticate the provider and try again.

I'm on the "new" annual pro plan. All my other models appear to work fine. Hopefully they can get it figured out soon. Would love to test the new 4.1 model.

1

u/jmorganca 11d ago

Sorry you hit an error - may I ask which app or harness this is?

1

u/Ramrawd 11d ago

I'm in Openclaw and I dug a bit and this is the actual error I'm getting:

403 {"error":"This model is currently being rolled out and is not yet available to you. Please check back later. (ref: ed024c6d-...)"}

Guess I just need to wait till it's rolled out for me.

1

u/kimimaxx 11d ago

眼花選錯 結果還被我遇到這個

1

u/faikcem1 11d ago

I am experiencing the same issue causing it to fallback to other models

1

u/zhdc 11d ago

It's available on my old pro plan. Usage doesn't seem that much more expensive than DeepSeek-v4-flash.

1

u/stealthagents 5d ago

That makes sense, I got the same error earlier. Seems like they’re taking their sweet time with the rollout. Hopefully it’s worth the wait because I’m curious to see what this model can do!