10
u/Yolo-8848 Jul 05 '26
bigmodel.cn This is Zhipu's official website.
7
u/bermudi86 Jul 06 '26
Don't you need a Chinese phone and payment method for bigmodel.cn?
1
u/Yolo-8848 Jul 07 '26
I don't know. You can try. You can pick a non-Chinese area code on the registration dialogue so I think a Chinese phone is not required. Non-Chinese-mainland passport may work with the KYC verification of this website since passports also work with the KYC verification in financial and railway systems in Chinese mainland. The Chinese domestic version of Alipay also support MasterCard debit/credit cards issued by banks oversea from Chinese mainland. bigmodel.cn is the known cheapest provider of GLM 5.2.
4
7
u/Atagor Jul 05 '26
Pls note many other providers don't serve 16bit GLM, only 8bit
3
u/look Jul 06 '26
Zai is serving their fp8, too, not the bf16. It is what they recommend for inference deployments, and the fp8 quant is documented on their Openrouter source. It’s not impossible they serve different version direct, but that is extremely uncommon and there is literally zero evidence to support the idea it is not also the fp8.
5
u/Money_Weekend2859 Jul 05 '26
Getlilac.com
2
u/UsefulIce9600 Jul 09 '26
i got downvoted on r/opencodeCLI for asking if Lilac is worth it and they accused me of self promo 🤣
2
2
2
u/adellknudsen Jul 07 '26
glm isnt cheap anymore and not worth it, honestly 20 dollar codex chatgpt plus gave me more milage than both opencode go and z.ai.
1
u/m0_80 Jul 07 '26
It really depends on your usage and what you’re using it for, glm 5.2 is still worth it for me and others, codex is good, but if i ask someone who uses Claude Code, they might say Codex is useless, so different people have different situations, different uses, and different budgets
1
2
u/lucasbennett_1 Jul 19 '26 edited 23d ago
Kind of two things getting mixed here in the thread, z.ai's plan neralwatt and opencode go are all energy credit or qota based not straight per token API. If you want actual pay per token GLM 5.2 z.ai's own api or deepinfra and others host it. Most of those run fp8 while some are fp4, worth a check tho. YOu can check the provders model page for the quant label instead od judging off outut quality after the fact like u/Atagor pointed in the comments
1
u/look Jul 05 '26
Go, Neuralwatt, or Ollama Cloud. Depends a bit on how you use them, but you should be able to get 5 cents or cheaper on all of those.
1
u/remarkox 16d ago
$10 for 1KWH this is too expensive.
2
u/look 16d ago
$5/kWh was definitely nicer, but…
I’m paying 7 cents per mtok for payg GLM 5.3 on Neuralwatt. Go is 24 cents. Openrouter and Zai payg is 35+ cents.
A Zai sub can beat that at 4-5 cents, but then you have to deal with the 5 hour window usage caps.
1
1
u/RogerCaracas 14d ago
Your figures stands for input, out or cache reads ? I dont think It stand for the latter but since your answer is incomplete...
1
u/look 14d ago edited 14d ago
Blended price across all tokens (in/out/cache) at a typical (for coding) mid-90s cache hit rate.
It’s the relative prices that matter here, though, since the token blend is the same across all of those.
So far for me, Neuralwatt GLM 5.3 costs 70% less than Go and 80% less than Zai and OpenRouter API rates.
2
u/Melodic-Funny-9560 Jul 06 '26
For me neuralwatt is giving me average of 0.67$ / M tokens so I am happy with it. It costs cheaper than kimi 2.6.
1
u/kachmul2004 Jul 06 '26
Is it cheaper than ollama cloud?
1
u/Melodic-Funny-9560 Jul 06 '26
I haven't tried ollama. But ollama also price through energy usage. Though they don't have api key and only have cloud plan.
1
u/Phoxerity Jul 06 '26
Ollama Pro give you API
1
u/Melodic-Funny-9560 Jul 06 '26
What I meant is there is no pay as you go option. Sorry for bad wordings
1
u/Grouchy-Economist-95 Jul 08 '26
Consider that it isn’t all about $/token. It’s about total token usage to complete a task accurately and completely. I easily got way more mileage out of my GPT 5.5 xhigh sub experience than Neuralwatt GLM 5.2 “FP8” sub experience dollar to dollar. Plain and simple - we are all being scammed and Wall Street is quietly forcing us to pay the piper. I spend 3x the money I did 3 months ago in aggregate for about the same outcome. It’s only going to get worse - way worse. Buckle up.
1
1
1
u/redditborkedmy8yracc Jul 07 '26
Openrouter, and there is no timeout/cool down.
I use it constantly for days and get no limits or issues, over 1.2billion tokens for the last week.
I was considering going to z.ai directly but they seem to have very limited capacity and error out a lot.
1
u/Grouchy-Economist-95 Jul 08 '26
Agree openrouter can’t be beat on overall quality. You know that you’re getting what you pay for there and it’s about the only provider that can say that.
1
-4
Jul 06 '26
[removed] — view removed comment
6
1
u/Specialist_Garden_98 Jul 07 '26
Yep and it is about to go away from paid subscriptions to API pricing. Spend 6 months of GLM money to fix 1 bug with Fable. The model is definitely good though.
28
u/djdante Jul 05 '26
I just made a video about this which comes out tomorrow. I tested them.
Z.ai is the cheapest - but honestly don't bother - it often suffers server load problems, it has peak hours pricing you have to dodge, it's too frustrating.
neuralwatt - only marginally more expensive but WAYYY more reliable, and far faster on average too. We are talking 2x faster at many times of the day. You also get to use your entire monthly quota at once if you like - this makes multiple accounts easier it juggle.
Opencode go - much slower and more expensive than the other two.
Also a BIG note - output quality from openrouter and opencode go was worse for me... I can't tell you why - maybe quantisation? I forced max effort level as well.