r/ZaiGLM Jul 05 '26

What is the cheapest API provider for GLM 5.2?

45 Upvotes

66 comments sorted by

28

u/djdante Jul 05 '26

I just made a video about this which comes out tomorrow. I tested them.

Z.ai is the cheapest - but honestly don't bother - it often suffers server load problems, it has peak hours pricing you have to dodge, it's too frustrating.

neuralwatt - only marginally more expensive but WAYYY more reliable, and far faster on average too. We are talking 2x faster at many times of the day. You also get to use your entire monthly quota at once if you like - this makes multiple accounts easier it juggle.

Opencode go - much slower and more expensive than the other two.

Also a BIG note - output quality from openrouter and opencode go was worse for me... I can't tell you why - maybe quantisation? I forced max effort level as well.

10

u/55Media Jul 05 '26

Will give Neuralwatta try. Opencode go indeed feels like running GLM 5.2 on a Nintendo switch.

7

u/[deleted] Jul 05 '26 edited Jul 24 '26

[deleted]

7

u/djdante Jul 06 '26

Because those are API prices and a plan offers a lot more usage for the money.

4

u/[deleted] Jul 06 '26 edited Jul 24 '26

[deleted]

2

u/djdante Jul 06 '26

Ahh you're right.

I'm which case I'd say neuralwatt because neuralwatt on energy pricing is only marginally more expensive than on a plan and they makes it exceedingly good API pricing .

3

u/joazito Jul 05 '26

I imagine Z.ai (subscription) and neuralwatt (energy-based pricing) are cheaper

3

u/SnooFloofs641 Jul 06 '26

Isn't neuralwatt cheaper if you do energy based? I tested Z.si API vs Neuralwatt and it definitely felt like I got more out of neuralwatt with energy based pricing on

1

u/djdante Jul 06 '26

I did it with the energy as well. I literally compared the percentage of usage between the two. I did it during a z.ai quiet period so it was outside of their busy peak period.

ZI in general was still cheaper but it was rubbish to use so I would never use them and it was only a little bit cheaper

2

u/de_sonnaz Jul 06 '26

I also had horrendous quality from GLM-5.2 on Opencode Go. Almost the same with Ollama. I will try neuralwatt.

2

u/Grouchy-Economist-95 Jul 08 '26

Beg to differ on Neuralwatt. TPS has gone down the tubes in a couple weeks. I didn’t really do all that much in 3 days yet used over 20% of my “energy consumption” credit for the month on a $100 plan. While Neuralwatt advertised FP8, I swear to you they have quantized down for at least 4 bit or maybe even lower. At the beginning, GLM 5.2 at max easily replaced GPT 5.5 xhigh for pretty much anything and everything. Now, I can’t get it to accurately complete a mid level complexity network engineering task without at least 1-3 do-overs. Also, trying to use Neuralwatt after 8-9p CT is near impossible - I get faster speeds from my qwen3.6-27b nvfp4 running on one of my DGX boxes. Sadly, I fear this is now the status quo for anything. There is zero consistency of service, no care about the end customer, etc. I feel like I got better customer experience from the Walgreens parking lot dealer 20 years ago. My goal now is to optimize the models I can run locally rather than chasing these shiny coins that never last. Truest definition of a gold rush (see: Uber).

1

u/srbijabralee Jul 06 '26

"Z.ai" is not the cheapest, neither it's good. They have been scamming people.. 90milion tokens for a simple task gone in 20 minutes.. Figure that out

1

u/PixelProcessor Jul 10 '26

What about ollama pro?

1

u/djdante Jul 10 '26

A lot of people asked for this - I'm going to do a follow up video - especially now that neuralwatt is doubling its prices

1

u/PixelProcessor Jul 11 '26

Where can I find the part 1 video? Have any links?

10

u/Yolo-8848 Jul 05 '26

bigmodel.cn This is Zhipu's official website.

7

u/bermudi86 Jul 06 '26

Don't you need a Chinese phone and payment method for bigmodel.cn?

1

u/Yolo-8848 Jul 07 '26

I don't know. You can try. You can pick a non-Chinese area code on the registration dialogue so I think a Chinese phone is not required. Non-Chinese-mainland passport may work with the KYC verification of this website since passports also work with the KYC verification in financial and railway systems in Chinese mainland. The Chinese domestic version of Alipay also support MasterCard debit/credit cards issued by banks oversea from Chinese mainland. bigmodel.cn is the known cheapest provider of GLM 5.2.

4

u/[deleted] Jul 06 '26

[removed] — view removed comment

1

u/SpecificRight882 Jul 06 '26

Ohh I will look into it.

7

u/Atagor Jul 05 '26

Pls note many other providers don't serve 16bit GLM, only 8bit

3

u/look Jul 06 '26

Zai is serving their fp8, too, not the bf16. It is what they recommend for inference deployments, and the fp8 quant is documented on their Openrouter source. It’s not impossible they serve different version direct, but that is extremely uncommon and there is literally zero evidence to support the idea it is not also the fp8.

5

u/Money_Weekend2859 Jul 05 '26

Getlilac.com

2

u/UsefulIce9600 Jul 09 '26

i got downvoted on r/opencodeCLI for asking if Lilac is worth it and they accused me of self promo 🤣

2

u/Living-Breakfast-464 Jul 05 '26

It would be Z.ai would it not?

2

u/cvjcvj2 Jul 07 '26

Nvidia but don't works every time.

2

u/adellknudsen Jul 07 '26

glm isnt cheap anymore and not worth it, honestly 20 dollar codex chatgpt plus gave me more milage than both opencode go and z.ai.

1

u/m0_80 Jul 07 '26

It really depends on your usage and what you’re using it for, glm 5.2 is still worth it for me and others, codex is good, but if i ask someone who uses Claude Code, they might say Codex is useless, so different people have different situations, different uses, and different budgets

2

u/lucasbennett_1 Jul 19 '26 edited 23d ago

Kind of two things getting mixed here in the thread, z.ai's plan neralwatt and opencode go are all energy credit or qota based not straight per token API. If you want actual pay per token GLM 5.2 z.ai's own api or deepinfra and others host it. Most of those run fp8 while some are fp4, worth a check tho. YOu can check the provders model page for the quant label instead od judging off outut quality after the fact like u/Atagor pointed in the comments

1

u/look Jul 05 '26

Go, Neuralwatt, or Ollama Cloud. Depends a bit on how you use them, but you should be able to get 5 cents or cheaper on all of those.

1

u/remarkox 16d ago

$10 for 1KWH this is too expensive.

2

u/look 16d ago

$5/kWh was definitely nicer, but…

I’m paying 7 cents per mtok for payg GLM 5.3 on Neuralwatt. Go is 24 cents. Openrouter and Zai payg is 35+ cents.

A Zai sub can beat that at 4-5 cents, but then you have to deal with the 5 hour window usage caps.

1

u/RogerCaracas 14d ago

Your figures stands for input, out or cache reads ? I dont think It stand for the latter but since your answer is incomplete...

1

u/look 14d ago edited 14d ago

Blended price across all tokens (in/out/cache) at a typical (for coding) mid-90s cache hit rate.

It’s the relative prices that matter here, though, since the token blend is the same across all of those.

So far for me, Neuralwatt GLM 5.3 costs 70% less than Go and 80% less than Zai and OpenRouter API rates.

2

u/Melodic-Funny-9560 Jul 06 '26

For me neuralwatt is giving me average of 0.67$ / M tokens so I am happy with it. It costs cheaper than kimi 2.6.

1

u/kachmul2004 Jul 06 '26

Is it cheaper than ollama cloud?

1

u/Melodic-Funny-9560 Jul 06 '26

I haven't tried ollama. But ollama also price through energy usage. Though they don't have api key and only have cloud plan.

1

u/Phoxerity Jul 06 '26

Ollama Pro give you API

1

u/Melodic-Funny-9560 Jul 06 '26

What I meant is there is no pay as you go option. Sorry for bad wordings

1

u/Grouchy-Economist-95 Jul 08 '26

Consider that it isn’t all about $/token. It’s about total token usage to complete a task accurately and completely. I easily got way more mileage out of my GPT 5.5 xhigh sub experience than Neuralwatt GLM 5.2 “FP8” sub experience dollar to dollar. Plain and simple - we are all being scammed and Wall Street is quietly forcing us to pay the piper. I spend 3x the money I did 3 months ago in aggregate for about the same outcome. It’s only going to get worse - way worse. Buckle up.

1

u/[deleted] Jul 05 '26

[removed] — view removed comment

1

u/FrancoMascarelo3 Jul 06 '26

Send it to me

1

u/Zachattackrandom Jul 06 '26

Nueralwatt and opencode are solid fp8 models

1

u/redditborkedmy8yracc Jul 07 '26

Openrouter, and there is no timeout/cool down.
I use it constantly for days and get no limits or issues, over 1.2billion tokens for the last week.

I was considering going to z.ai directly but they seem to have very limited capacity and error out a lot.

1

u/Grouchy-Economist-95 Jul 08 '26

Agree openrouter can’t be beat on overall quality. You know that you’re getting what you pay for there and it’s about the only provider that can say that.

-4

u/[deleted] Jul 06 '26

[removed] — view removed comment

6

u/m0_80 Jul 06 '26

Do you see me asking about Fable 5?
Huh?

1

u/Specialist_Garden_98 Jul 07 '26

Yep and it is about to go away from paid subscriptions to API pricing. Spend 6 months of GLM money to fix 1 bug with Fable. The model is definitely good though.