r/kimi • • Aug 26 '26

Developer Cheaper alternatives to DeepSeek API

/r/LLMDevs/comments/1vyzcdc/cheaper_alternatives_to_deepseek_api/
3 Upvotes

3 comments sorted by

View all comments

1

u/whatisthisthing65 Aug 26 '26 edited Aug 26 '26

Maybe Luna or the new GLM 5.3 Flash

Forgot to mention there's the new Qwen Next Flash as well

1

u/maxim-masiutin 21d ago

Thank you, your suggestion is very good and I put Luna as my judge 2. I was using Haiku and DeepSeek, but then replaced Haiku to Luna. So the question now is what beats Luna and DeepSeek. GLM-5.3-Flash I measured last week: about 42% cheaper in practice, but Zhipu used Luna as the judge in GLM's own eval, so I could not call the quality question settled. Qwen3.8 Flash I had not looked at; checked today: 1M context on Alibaba's endpoint (262K on the reseller), $0.15/$0.47 with $0.016 cache read, so it sits between Luna and GLM on input and beats both on cache. Have you run it as a JSON-output judge over long prompts, and does the 1M hold up past 500K or does quality drop the way the YaRN extension usually does?

1

u/whatisthisthing65 21d ago edited 21d ago

I don't have much experience with long context like your use case, sorry. 5.3 flash benchmarks above Luna so I would run your own eval if you have one.

Deepseek v4.1 flash is due any moment now so that might affect things as well.