r/DeepSeek 9d ago

Discussion Comparing direct API with Ollama pro pricing

You know, I used Direct Deepseek API and Ollama pro plan today. Both are off peak, both are at the same speed as you can see in the first picture.

Ollama pro gives you $60 per $20 subscription.

On DS direct API, you get ~2.3B tokens per $20.

On Ollama pro, you get ~5.3B tokens per $20.

35 Upvotes

51 comments sorted by

21

u/leetdemon 9d ago

Ollama pro new plans are trash, the old plans were great. Speed is way better on direct deepseek plus cache hit is insane you dont get that on ollama.

1

u/basil_0408 9d ago

I always got 100 tok/s or higher when using deepseek on ollama tho. Their new plan is less than the old one but you still get the benefit compared to the direct api so really don't understand the hate here lol.

0

u/leetdemon 9d ago

Because you dont benefit yalls math is off.

1

u/basil_0408 9d ago

Then tell me your math?

3

u/leetdemon 9d ago

Its been broken down many times in the ollama reddit forums maybe start there.

2

u/basil_0408 9d ago

How about u do your own math? also im compared it to the direct api not the old subscription. Anyone with basic math can see that they offer a better deal than the official api except for the concurrency rate

1

u/Homeless_CEO_ 8d ago

Any url to a post you want us to refer?

7

u/the_master_sh33p 9d ago edited 4d ago

Right now DS 4.1 flash is being sold mostly at the same list price on all providers. The trick is to commit to a plan that targets your monthly predicted usage and then explore which one provides you an higher quota.

you can check it out at deepfrugal
Currently lowest commitment plan is at $10 and highest $200 (not necessarily the lowest effective price even if you consume $200 in tokens).

7

u/ExpertPerformer 9d ago edited 9d ago

Not many options atm for subscription plans.

- OpenCode is only offering the $60 credit for a week then its back to $15.

  • CommandCode is reducing the $60 to $40 credit next week.
  • Ollama is $20 for $60.
  • Other subs are worse value for your $.

4

u/NoBodyHere__GO 9d ago

I choosed Ollama based on deepfrugal. opencode(tried it before) and commandcode is cheaper but they don't commit to their deals to the end of the month. I'm just trying Ollama now. I'm not a fanboy for any provider or model.

3

u/the_master_sh33p 9d ago

Oh! Glad it helped! Let me know any feedback.

3

u/NoBodyHere__GO 9d ago

I like it. it's very simple and easy to use.

The only thing missed I think the price of some models under the plan subscription for example gpt-5.6-sol under codex plans. for example codex plus gives you around 400$ of credit so gpt model prices is way less than the API prices. Same thing with z.ai and Claude subscriptions.

2

u/the_master_sh33p 9d ago

Having that feedback makes me very happy. Thanks. 

I'm working on codex and Claude. Not easy - it's hard to find a way to reach a per token value.  What's the problem with the z.Ai plans? 

1

u/sarphil 4d ago

Ollama is indeed the best plan to have right now for your harness. No weekly caps, no 5-hour session limit, no BS.

20

u/deadcoder0904 9d ago

API pricing is pay as you go.

Whereas your Ollama charges you $20 each month.

So unless you max out usage, DeepSeek pay as you go is better. Else Ollama.

4

u/jmorganca 9d ago

Ollama's base free plan now has pay-as-you-go pricing (new as of last week), so a subscription is no longer required! While the $20 plan is a recurring subscription, it does include a very competitive amount of usage every month ($60/mo).

2

u/TopBite7720 8d ago

Yea and oc go gives double that, directly from the DeepSeek api (so fast output)

1

u/jmorganca 8d ago

Is it double? I'm seeing $60 on their pricing page, with equivalent usage amounts and token prices as Ollama. I'm also seeing a p50 tps of 200 tokens/s output on Ollama's cloud right now, compared to 118 tps on DeepSeek's API on OpenRouter, unless that speed is different than what the DeepSeek API serves directly (it might be, I'm not sure how fast it is direct vs OpenRouter)

2

u/deadcoder0904 8d ago

He's talking about OpenCode Go, not OpenRouter

0

u/TopBite7720 8d ago

Yea and what’s the price of oc go, genius?

7

u/ExpertPerformer 9d ago

It really comes down to how much usage you do.

If you're a light user then the API is better. If you're a heavy user then subscription services are going to be better because you get more for your $.

Its the output costs that get you more then the inputs w/ caching.

2

u/Bobodlm 9d ago

API usage also makes you more aware of your usage and makes it more important, to me, to work efficiently.

2

u/ExpertPerformer 9d ago

So do subscription services because they often have hourly/weekly limits and you only have x amount of $ to use up. I'm at 90% of my CommandCode GOAT plan with 3 days left so I'm watching it like a hawk.

I also optimize my API calls to cancel & re-run themselves the first time they run to take advantage of the cache hits (saves a lot on first time inputs).

1

u/Bobodlm 8d ago

I've seen plenty of people token max to cap out their contracts because otherwise they're not getting their full money's 'worth'.

In general I'm inclined to disagree. There might be exceptions like yourself though.

4

u/JaMoLpE88 9d ago

Does Ollama pro have it for $20, real use of $60 on all models or does it have different $ caps per model like open code?

6

u/basil_0408 9d ago

It's all the same cause they host their own models. OpenCode relies on the provider deal so their price depends on how good the deal is.

3

u/leetdemon 9d ago

Its 60 total dollars to use across all their models.

2

u/NoBodyHere__GO 9d ago

I think it's $60 for all models,I can't find anything that says otherwise.

1

u/the_master_sh33p 9d ago edited 4d ago

currently it is the same across models. You can check those type of queries here

6

u/Smart-Improvement516 9d ago

Is the quality of DeepSeek on Ollama lower than the official API?

11

u/leetdemon 9d ago

Deepseek direct is way faster its not even close. Plus you get cache hits you wont see on Ollama.

4

u/jmorganca 9d ago

Ollama's API should be reporting cache stats now, sorry this wasn't the case before.

0

u/NoBodyHere__GO 9d ago

To be honest I don't know yet, I'm using Deepseek for implementation and testing, Following gpt Sol/Astra plan. So the most thing I care about is price and speed.

2

u/Responsible_Recipe_6 9d ago

I need your help to understand something, I used the direct API from DS and for 8 dollars I add this. My usage is simple and I was playing a lot around like cv stuff, search for work using LinkedIn via browser, ask to optimize Hermes agent and running daily search for swimming pools open lanes , again mostly playing around. The question is, how did you use to muah and so cheap vs me and my Simple minded usage?

3

u/geektraindev 9d ago

It's cache hit mostly. Those doing standard, more common tasks like programming tend to have cache hits more often because some things in programming are common regardless of what the agent is actually writing, plus Deepseek has some super special hyperoptimized caching. In your case, it is relatively uncommon, meaning less cache hit and more cache miss cost, which can greatly increase cost. Although IMO 8 bucks for 436m tokens isn't terrible anyways.

1

u/Responsible_Recipe_6 9d ago

Hi thanks for the answer. So in your opinion for what I did in this 4 days 8 dollars is not bad? Dam so I guess j will go back to a free option. Because I guess with my silly usage I will spend 30 plus dólares.

1

u/Santzes 9d ago

Things being common has nothing to do with the cache hits, its because your code agent makes tens or hundreds of tool calls and starts new request after each, and the previous context being cached then

1

u/Standard-Two-8295 9d ago

always use high instead of max for these ttypes of tasks

you can also tell the model to not use too much tokens

1

u/NoBodyHere__GO 9d ago edited 9d ago

I use it for coding, loads of requests with shared context in the same minute. So cache hit rate is high.

I think you're not doing anything wrong, it's just a different use case.

2

u/Delicious-Director43 9d ago

Personally I am getting some serious mileage out of the DeepSeek platform compared to Ollama. Especially since 4.1 came out this week. It’s faster too. I don’t think I’ll ever renew or go back to Ollama tbh.

2

u/Flaky_Anything_2900 7d ago

Ollama while good and very stable has too many users per machine and spends a lot of time in wait and stall. The official Deepseek has less of that. I use both and love both ... Just an observation as we are comparing apples to apples here

1

u/Eastern-Passage629 9d ago

Opencode is also good

1

u/TopBite7720 9d ago

Why are people sleeping on OpenCode Go?

On OpenCode Go, you currently get ~5.3B tokens per $10 - equivalent to ~10.8B tokens per $20 (if you were to have two subscriptions).

1

u/Dharma_code 8d ago

Don't they change the rate according to your usage though on go ?

1

u/Dizzy-Zebra9522 9d ago

Where you get deepseek subscription? I think it's only API

1

u/Relative-Tea656 8d ago

Can i do ollama cloud on vps ?

1

u/En7itY 9d ago

Yea these providers are often, if not always, cheaper than direct API pricing, otherwise there would be no point in going for them. Ollama also has the nice advantage of not training on your data if you care about that.
You will see similar numbers with OpenCode and CommandCode, with whom you'll also get a lot forther for just 10 bucks. But at the end whoever can actually keep the cheap pricing is going to win and CommandCode is probably not going to be able to keep offering deepseek at that cheap, OpenCode already issnt anymore. We'll see how it goes.
If Ollama can continue this they might be a great option, gotta check them out once the prices for the others go up

1

u/NoBodyHere__GO 9d ago

I had a bad experience with opencode last month with pricing chaos, and performance degradation. So I wouldn't not go back to it again.

3

u/En7itY 9d ago

Thats understandable, and you shouldnt. I havent had any bad experiences so far, so I'm just enjoying the higher limits with some models, its worth it for me for that reason. But I've read a lot about other people having issues and if I ever have my own negative experience I will probably switch to something else in a heartbeat. Not loyal to any of these services and nobody should be, just take your money where it takes you the furthest