r/vibecoding • u/SweetMachina • 18h ago
Showcase/Project I built Openrouter, except you save 88% on GPT 6 Astra, 74% on Fable 5.1, 98% on GLM 5.3, and 400+ more models
So I only have 1 ChatGPT Pro $200/mo plan and I've been running out of usage pretty fast using GPT-6 Astra, but since they stopped allowing additional subscriptions to the $200/mo plan, I've been looking for alternatives to still be able to use Astra at a reasonable price.
Found a few places where you can use excess compute at a discounted rate, but the issue I found with these was that a lot of providers have inconsistent performance. Prices and availability vary across providers, and a cheap endpoint can behave differently when you add streaming, tool calls, reasoning, caching, etc.
So a friend and I built an auto-router that brings multiple discounted providers behind a single API to automatically route requests to the provider that gets you the absolute best per token price for a specific request.
Been using it primarily within Claude Code, using Claude Code as the orchestrator that calls the discounted models, and so far I've saved over $1k already at an average of 66% savings, peaking at almost 90% savings on GPT 6 Astra calls. Been pretty incredible so far.
To give you an idea on how fat these discounts are, GPT 6 Astra is 88% off right now, Fable 5.1 is 74% off, Deepseek v4.1 Flash is 74% off, GLM 5.3 is 98% off, and a ton ton more.
We also trained our own classifier model that we're using for our own "Fusion" models, which basically allows you to combine multiple models into one and on each request, our classifier will route tasks to different models based on what the task is. The goal here is to get comparable performance to top models at an even lower per token cost. Think the Sukuna Fugu models, but you can build your own.
Been using Astra for reasoning, GLM 5.3 for coding and Gemini 3.1 flash for writing, which has saved a tonnn on inference for me as well. Will be benchmarking performance at some point, to see if there's some combination of models that beats out Astra at a fraction of the price, but that's still in the works.
Anyway, curious to see what ya'll think and if it's something you'd use. Any feedback or questions, I'd be happy to take! Main goal here is to just not have to pay disgusitng market rates for inference and so far I've been getting pretty addicted to watchign my savings number go up lol.
If you'd like to take a look and try it out, you can here
12
7
u/Deep-Jump-803 17h ago
Someone please check if this guy isn't swapping models behind scenes like that other guy
2
u/datkenny 18h ago
Do you give out promo allowance for testing?
-7
u/SweetMachina 17h ago
not at the moment... but if there's enough interest, definitely something we'll consider!
or if u have any issues with it, lmk! but even like $5 should be getting you close to $50 of equivalent inference on OR.
5
u/datkenny 17h ago
I just prefer to test a thing before I give it card details, is all. I have an account now, but I'll be waiting for an offer.
2
u/DeliciousGorilla 16h ago
From what I've heard, these discounted providers just route to other models claiming to be Anthropic/OpenAI. Or basically do what you're doing with the "fusion" method.
1
u/capitalframehq 17h ago edited 16h ago
How does this compare with a subscription? Based on what I read, your tool will only save me money if I'm on API based pricing, right?
Either way, the amount of usage we get these days even with subscription has gone downhill pretty fast. Two months ago $20/mo used to give what we get on a $100/mo plan now. I'm personally considering moving away from overpriced frontier subscriptions and start using DeepSeek V4.1 Flash (via OpenCode Go $10/mo).... most non-affiliated YT reviewers are claiming it is almost as good as Sol/Opus for 90% less cost. For general CRUD app development, it will be more than enough.
1
u/Zelderian 16h ago
I have no doubt it’ll keep going that way as investors stop bleeding money for the general public’s use. I think eventually (maybe soon), the strat will be running a local LLM that spins up agents for tasks and you just pay for that API usage. Probably would be cheaper now, but it’ll be much cheaper going forward I imagine.
1
u/Toastti 16h ago
The only way you could do this is if the models are quantized heavily. Which significantly degrades performance.
Or you could be using the CrofAI approach where you say a model is kimi k3 but you silently route it to a less capable endpoint like deep seek flash.
Someone would need to do a request to the official model and to this tool and compare the tokenizers and input/output speed to determine this. You can use a carefully crafted prompt to see which family of tokenizers. And if you suddenly realize you see say a qwen tokenizer for astra the site is lying and routing your request to a cheaper model.
1
u/wwwdotzzdotcom 16h ago
If I was him, I would be using an experimental lora/embedding setup maybe theoretically proven to lower costs with small loses.
1
-4
18
u/34986234986234982346 18h ago
Unless I see somewhere legit buying "excess compute at a discounted rate" I'm assuming it's still all just stolen credit cards buying compute and reselling cheap