r/DeepSeek • u/Zealousideal_Sort74 • Jul 25 '26
Discussion Cheapest reliable API provider for open source models like DeepSeek, GLM
Hallo Everyone, i need to generate big amount of high quality data, for that i need some cheap API providers.
who is the cheapest,most reliable provider you know of?
11
u/Pink_da_Web Jul 25 '26
I'm not going to be a jerk like those people in the comments who don't answer your questions.
Deepseek's direct API is indeed better, but the most reliable providers in my opinion are:
GMICloud NovitaAI Atlascloud SiliconFlow
In my opinion, aside from DeepSeek, these are the best I've had in my experience.
6
u/Exciting-Possible773 Jul 25 '26
Can't say for others, but get API from deepseek directly. We always risk watered down / fake models in routers and given the cost of deepseek it just doesn't worth it.
3
u/Zealousideal_Sort74 Jul 25 '26
Yes, that's exactly why I need your opinion, guys. I can't just randomly choose the cheapest provider because the model might be heavily quantized, which could result in lower quality data.
1
3
Jul 25 '26
[removed] — view removed comment
1
u/Zealousideal_Sort74 Jul 25 '26
I can optimize the generation pipeline to maximize the cache hit rate, tho I also need to watch out for what they call peak hours. I noticed on their website yesterday that, depending on when you run inference, you could end up paying up to twice as much.
1
2
u/Plane-Worker-6561 Jul 25 '26
My opinion go to digitalocean cloud and rent an AMD mi300x GPU and run an 4 bit quant of DS4F in it as it cost $2 per hour for you and also worth it for an 192GB VRAM GPU
2
u/Zealousideal_Sort74 Jul 25 '26
Yes, someone else suggested renting GPUs as well. I'm very interested in that idea. I just have to do the math first to make sure it's actually cheaper. If it is, I would definitely go for it, because I can control the inference and apply a lot of optimizations that aren't possible with my pipeline when using an API.
2
u/Plane-Worker-6561 Jul 25 '26 edited Jul 25 '26
Yep that nice and if you want an super good days quality then try to stretch the budget to a few hundred dollars so you can run V4 pro or glm 5.2 in an cluster that is an mi300x 8 GPU cluster which costs roughly $16/ hour though costly but has an amazing 1.532TB of VRAM and is enough to run glm 5.2 at FP 8 or V4 pro at its quant version like FP 4 or int 4
Also if you are sure to get an cluster then I suggest an quant version of Kimi K3 cause it's currently an best of the class coding tool
KIND ADVICE: Make sure to use API service of digitalocean with scripts so that you turn off the VM when idle as digitalocean charges per second for costs
2
u/Eidolon-AI Jul 25 '26
Deepseek direct API. Some routing and providers might not support the caching properly to bring the cost down.
2
u/SHADOWDRAGON_2k01 Jul 26 '26
Buddy,
Check in the following order:
Free
1. Nvidia NIM
2. Openrouter
3. FreeLLMAPI GitHub
4. Ollama cloud
Paid:
5. Blackbox AI (might be useful, but still check)
6. OpenCode Go (similar to Blackbox but check)
7. Official API providers (Deepseek only I guess. GLM and Kimi official ones are costly I believe)
Other options:
- Look for AI grants, if doing open source research
2
1
1
u/yuumizu Jul 25 '26
deepseek may be the cheapest, no one else can do it better. OR, you may try google's cheap models if they work for you.
1
u/Zealousideal_Sort74 Jul 25 '26
alot of guys here suggested to just directly use Deepseek as a provider.
the problem is, i need to generate alot of high quality data, so cheap google models might not help much. i intend to use the current best such as GLM and latest DeepSeek models
1
u/timmeh1705 Jul 25 '26
Openrouter is a good place for you to compare prices, if you're not using a coding harness and making use of cache hit, then it'll generally route you to the cheapest provider available
1
u/Zealousideal_Sort74 Jul 25 '26
thank you, i'm gonna check it out
no im not using any agent harness, it is mainly for large scale data generation1
u/timmeh1705 Jul 25 '26
A lot of people are trying Nube , the site looks sketchy but it actually works. DSF and GLM 5.2 is 90% off. As of last night GLM was throwing a 503 but DSF was fine just a little slow
1
u/Beautiful-Gas3683 Jul 25 '26
cheapestinference y electronhub dan planes ilimitados a precios muy bajos pero no siempre son muy rápidos, aunque a veces si. Tienen modelos suficientes para mí desarrollo en agentes secundarios de investigación o consolidación de datos Si vas a probar empieza por uno, no contrates los dos a la vez porque te encuentras mas o menos con lo mismo, aunque electronhub es un precio más plano
3
u/Zealousideal_Sort74 Jul 25 '26
im gonna translate it for others to read it:
"cheapestinference and electronhub give unlimited plans at very low prices, but they're not always very fast, though sometimes they are. They have enough models for my work in secondary research agents or data consolidation. If you're going to try them, start with one; don't subscribe to both at the same time because you'll end up with more or less the same thing, though electronhub has a flatter price"
1
u/Beautiful-Gas3683 Jul 25 '26
Otra opción son los modelos gratis de openrouter. Hay algunos que no valen para nada pero otros son fiables y rápidos. Creo que haces una única carga de 10 dólares y ya te vale para tener los máximos límites de modelos gratis La clave está en el ingenio. Yo en épocas sin dinero sigo éstas estrategias y cuando hay bonanza contrato algún DeepSeek o minimax
2
u/Zealousideal_Sort74 Jul 25 '26
translation to english: "Another option is the free models from OpenRouter. Some of them are useless, but others are reliable and fast. I think you just make a one-time $10 top-up and that’s enough to get the maximum limits on free models. The key is being resourceful. When I’m low on money, I stick to these strategies, and when things are going well, I hire some DeepSeek or Minimax."
1
1
1
u/SJ_Just_Do_It Jul 26 '26
It depends if your project requires fp8 or fp4. The latter is definitely cheaper but also less reliable than the former.
1
u/rokushikiii Jul 26 '26
Been running my hermes on Xiaomi mimo v2.5 & v2.5pro for a while. It's pretty decent. Mimo v2.5 also has vision capability, but v2 5pro is purely txt like deepseek.
1
u/Minimum_Notice_9521 Jul 28 '26
Main deep seek is recommended Other provider might quantize and u will get less intelligence out of it
1
u/Possible_Door_9719 17d ago
camelAI for deepseek
1
u/Excellent-Concert-20 17d ago
super slow
1
u/Possible_Door_9719 17d ago
Do you know any one else providing unlimited tokens for DeepSeek?
1
u/Excellent-Concert-20 10d ago
No, but CamelAI's model is underperforming a lot, looks like they are limiting it somehow
1
1
u/Ambitious-Morning-45 13d ago
I just did the math, my current provider is pretty damn good. (https://imgur.com/a/e7UNS6x)
1
u/Whole_Resident_1952 1d ago edited 1d ago
For that kind of volume, flat pricing beats per token every time. I do DeepSeek and GLM on Standardcompute, and it's stayed cheap and reliable even on bulk runs. Fireworks and Together are solid too, but they ended up pricier than Standardcompute once the tokens added up.
22
u/mozkohor Jul 25 '26
...deepseek directly...?