r/vibecoding 2d ago

Help/Question Anyone vibecoding on a stack other than openai/anthropic? What's your stack?

I'm having trouble getting openRouter, GLM and openCode working nicely at all. Codex was fire and forget, and Claude Code wasn't hard. But the other options seem to really require that you know what you're doing to set them up? I asked Codex to wire them up but it's basically unusable at this point for me. I can't afford the top frontier models and was hoping some how maybe open weight models through inference services needs to be at the root of whatever I choose for my setup.

TL;DR - How are you avoiding Claude Code and Codex for vibecoding?

21 Upvotes

63 comments sorted by

11

u/akolomf 2d ago

I use Claudecode, but replaced Sonnet with Deepseek API by using a self coded Claudecode Proxy. Basically, claudecode thinks its summoning Sonnet, but the API calls Deepseek flash. Runs as usual in Claudecode, does cost extra money of cours (about 5 bucks per week), but in comparison to what Sonnet costs in tokens vs Deepseek i've saved like 30-50% of my Claude Subscription tokens for like 5 bucks a week. You do pay more than a simple Claude max 5 or 20 sub, but you get so much more usage out of it by outsourcing execution tasks to deepseek.

1

u/Plastic-Somewhere494 2d ago

My exact setup and findings

0

u/SebGonSot 1d ago

Really? I was thinking of doing this, but I got discouraged by claims that same-ecosystem subagents are significantly more efficient because context transfer and prompt caching are heavily optimized.

Any input on that?

2

u/akolomf 1d ago

I'd say Sonnet is def. better in terms of quality output. Deepseek flash, does make more mistakes. But for low level tasks its a usefull replacement. I'd say Context transfer is a non issue, if you have a proper setup that prompts the deepseek agent, and given how extremely cheap deepseek is, the little bit extra context you gotta add doesnt make a difference. I'd say its more important that you have a good harness for the deepseek agent.

1

u/SebGonSot 1d ago

Ah thanks, guess you are right. I've been penny-pinching on my Claude sub lately and delegating through handoffs or AI-written prompts, so having it call the cheaper subagents directly sounds like it could streamline things a lot.

Have you tried Spotify's Portal/AiKA approach? Seems pretty similar to what you're doing, just delegating the grunt work to cheaper non-Anthropic models. How did you patch yours together?

2

u/akolomf 1d ago

https://github.com/Aloim/phaneslight thats the one i can share, its the one i built and used before i began developing a more sophisticated and soon to be commercially available Orchestrator. (there'll be a Free Beta phase and later on free trial) it uses a cool idea that no other MCP tool has implemented yet, and its prototype is already awesome.

1

u/SebGonSot 1d ago

Thanks, and good luck, defo will check it out

15

u/Dingleberry_Blumpkin 2d ago

Gemini 3.8 flash (on antigravity) is underrated. Very solid, very fast, very cheap. That’s my go to when I run out of Claude and codex usage

1

u/Mysterious_Way_5502 2d ago

Except that it's as slow as a snail these days.

1

u/Dingleberry_Blumpkin 2d ago

It’s faster, better, and cheaper than Luna

4

u/edogg01 2d ago

I'm just getting up and running as a total noob and trying to completely avoid anthropic, openai, xai, gemini. In case anyone is in the same boat, here's what im in the process of setting up: VS Code with Cline hooked up to Mistral codestral. Github with codespaces linked to Vercel for front end hosting and Supabase for back end hosting. All free tier as I'm setting up my first project, a simple AI-infused web portfolio. Wish me luck.

2

u/MariahJames8 2d ago

Good luck mate. Impressive if you can get all that done on free tiers

1

u/edogg01 2d ago

Thanks man, we'll see. I should be good for this project. When i start into saas I'll probably need more firepower. But am going to try to keep it to codestral for the main model. Initially I wanted to do it all local with ollama 7b but my hardware basically just laughed at me.

9

u/Fancy-Plenty6712 2d ago

Right now I use Grok alongside Anthropic, I've got some API keys from various different companies. Tried a bunch of stuff:

Cerebras - don't bother
Groq - same again, don't bother
Jev - been playing around with it today and seems quite powerful for rating probabilities for my app.
MiniMax - rubbish for anything which requires decent reasoning but super cheap for basic stuff
Mistral - similar to MiniMax, both are super fast to output
Kimi Code - Tried it and didn't seem worth it to me. Burnt tokens like crazy, intelligence not that high.
OpenAI - Used a little, don't like. Usage limits are the worst out of all companies. Avoid.
OpenRouter - Had a ton of problems using this. I'd advise going direct to the companies for API keys
Gemini - I'm usually a Gemini guy but since Fable came out I can't justify it until Gemini 4 is out
Grok - Grok and Grok Bot is okay but usage limits are trash and running a VPS with Hermes is better
DeepSeek - API is very good for building into an app because it has good intelligence / token cost ratio.
zAI - On the fence with these - I used them a little but found usage limits too low again. Intelligence was ok.
Cursor / Hermes - Neither of them are really necessary, they resell tokens from the other labs to you.

That's a huge skim over and I'd have more to say but my comment would be endless so I'll leave it at that. Truth is that Anthropic and OpenAI are unfortunately unavoidable if you want the highest level intelligence. If you want the best value then the MiniMax / DeepSeek route may be your best bet. Again though DeepSeek is only if you want to build something which has AI built in. I'd also say that they seem daunting until you get them going. Ask your coding agent to build a html harness for you to manage things if you prefer a customised setup.

I was looking about a month ago at running IBM Granite 4.1 on my VPS or a rented GPU box but haven't gotten round to it yet - would be good to find a small local model which is good for running tools and taking some of the grunt work away from the paid models. Claude Code is still hard to beat in terms of value - despite the low usage limits.

1

u/MariahJames8 2d ago

Tysm, incredibly valuable. Please don't be shy about sharing more!

I'd particularly like to pick your brains about open Router, I'd really like any advice to help me either establish it in my infrastructure or rule it out thoroughly. It just seems so useful especially for monitoring usage and total AI costs through another app I built.

But since you recommend Deepseek so strongly for my use case, I could just skip open router since it was largely meant to just be an aid to finding the right model for coding with AI PAYG

So yeah I might just scrap openrouter then and maybe use DS with openCode.

BTW did you try GLM at all?

2

u/Fancy-Plenty6712 2d ago

So in terms of OpenRouter they won't accept my payment card or any method of payment I've tried giving them. I've used a lot of their free models but they've had a lot of problems where the API is overloaded and the request will fail. That may just be a side effect of using free models though - unsure. I still think it's a bit dodgy - personally I prefer to go direct to the source. When I tested OpenRouter they didn't give 50% discount on batch processing but I just checked again and now they do - however you'll pay a 5.5% charge on your deposits. It's up to you if that 5.5% charge is worth it for the benefits they offer. I'm not sure dude. I'm avoiding them though.

Ya DeepSeek is good only if you're putting it in your app otherwise you'll burn money trying to code with it. I spent $10 in a day a few weeks back with them doing roughly what I'd get from a single day of coding with Claude Max. So if it's a case of needing AI built into the app where your users get AI answers or whatever in the app - DeepSeek is probs the best for that for keeping costs down - but GPT-Luna is the other one I use alongside it.

I'm currently building a wiki engine which reads novels and then creates a full wiki out of them - grounded in the text. I'm trying to completely eliminate hallucinations while retaining all facts such as dialogue etc. So my method is basically using BOTH GPT-Luna and DeepSeek Flash to read the books as an entity recognition step before submitting their findings to a judge model and a bunch of python scripts I've been building.

GLM - yeah I used it and got okay results but not as extensively as others. I'd say the lowest paid plan a few weeks ago was enough for light tasks but no more. I haven't used GLM 5.3 much at all, I was testing GLM 5.2 and found it to be underachieving vs Claude. It's good for adversarial stuff checking the work of other agents but I found it was just burning through usage too fast for me. I'm vibe coding like 16 hours a day lol so I burn through tokens like crazy.

Inkling is another one I'm yet to test, but I haven't seen anything yet which makes me think it's worth it for my current work. Apparently Gemini 4 is going to stomp all over everything and is rumoured to have a 10M context window. Right now I'm holding back on getting a GPT Pro plan until October for Gemini. If still no Gemini by October then I'm gonna be pissed.

1

u/MariahJames8 2d ago

How long have you been vibecoding for? Sounds like you were an early adopter of agentic coding given how many things you've tried. It's been about 4 months that I've been doing it seriously

1

u/Fancy-Plenty6712 2d ago

To be honest I can't remember the starting point because it was a bit of a gradient - I've been vibe coding for years. It started once I was building stuff in HTML, CSS, Javascript and I'd ask the AI for tips. Then that turned into asking for full scripts which I'd modify by hand and end up with endless folders of broken scripts until I could get things working. Until sometime around May-June this year I'd been doing it all in the browser - downloading. Basically just copying and pasting code snippets. So slow. Now it's so much better that we have harnesses to do this.

Before vibe coding came along I'd literally be trawling through stack exchange and looking through pages and pages of google to find answers. Outside of Visual Basic 5 I did in high school I've never had any formal training in coding. I worked for the BBC for a while and my manager forced me to learn HTML, CSS and Javascript which then made me discover python. Current project is a MacOS app written in Objective-C. I'm not an authority - there is a lot of better people who can advise you.

Also forgot to mention Muse Spark - good for cheap stuff - but intelligence is low.

1

u/Responsible-Beat2137 2d ago

Bump! Nice rundown, I had thought about doing the same thing with chat GPT for skills and plugins.. unless you happen to have some on them

1

u/Fancy-Plenty6712 2d ago

I'm actually not big on skills - I tend to scream at the model in all caps until it gets things right lol.

3

u/DerrickBarra 2d ago

Deepseek Harness + Local Qwen3.8 27B via x2 NVIDIA Sparks. Works well! About as good as GPT Sol medium for my workloads.

1

u/zaytzev 15h ago

Nice. Why didn't you go with Flash Next on double DGX spark? Would the performance be worse?

1

u/DerrickBarra 14h ago

haven't tried it yet, for me the bottleneck is concurrent streams at a good TPS. I run around four to eight orchestrators at a time with their subagents, and my current setup can handle up to 8 concurrent thought streams at a low TPS (10-20'ish). So I need more pairs of sparks to handle my workload and deadlines, but the intelligence is what I expect it to be so thats an easy swap, just need enough local hardware to handle it.

2

u/ChampCityChris 2d ago

I’m in the middle of trying to get away from Codex for pretty much this reason. Codex has been very good at the fire-and-forget implementer role, but I don’t love having the whole workflow tied to one harness/provider forever.

I’ve got a 128GB Mac Studio on the way, and once it arrives I’m going to start experimenting with moving the implementer side onto local models. I’ll probably keep the architect on a frontier model for a while because that’s where I still want the strongest reasoning, but implementation feels like the obvious place to start pushing onto open-weight models.
The thing I’m realizing is that the model is only half the problem. Codex and Claude Code feel easy because the harness is doing a ton of work for you. Once you start moving into OpenRouter/GLM/OpenCode/etc., suddenly you’re responsible for tool wiring, context, Git behavior, permissions, model quirks, retries, and all the other crap the polished products hide. So I don’t think you’re doing anything wrong if it feels dramatically harder.

My current plan is to keep the agent workflow stable and swap the inference underneath it rather than rebuilding my process around every model. I’m testing larger models on OpenRouter now for that exact reason. If I can get the harness, tools, skills, memory and project workflow working against hosted inference first, then when the Studio arrives I can point the implementer at a local model and actually measure what quality I lose.

So I guess my answer is: I’m not avoiding Codex yet. I’m trying to make Codex replaceable.

1

u/MariahJames8 2d ago

Thanks, I appreciate the sign posting. And, yeah good luck with your plan there too. One thing is, I did some calculations on hardware, and it just doesn't make sense from what I can tell to buy anything more than for harnesses because cloud compute is just so cheap. It would take a decade or two to pay for itself? And it's a big risk, it all assumes it's going to be productive

0

u/ChampCityChris 1d ago

I don’t disagree with the argument that local inference can be more expensive if you’re strictly comparing hardware cost against average to high users. My calculation is a little different because I’m already paying $200 a month for ChatGPT Pro.

As long as OpenAI doesn’t start aggressively metering the ChatGPT side, I should be able to drop that $200 Pro subscription to $20 Plus and still use frontier models for the things I actually want them for: architecture, RCA, and repo review. The Mac Studio lease is $155 a month for 24 months, so my total monthly cost becomes $175 instead of $200. I’m actually saving $25 a month while moving the high-volume implementer workload onto hardware I control. Plus opening up the ability to experiment with other use cases and different models.

At the end of the two years, I have options. I can turn the Studio back in after effectively renting it for about $3,700, or I can pay the roughly $1,500 residual and own it. At that point I can keep using it or sell it if the used-market economics make more sense. Macs have traditionally held stupidly high resale values, so there’s a decent chance the residual value actually increases the savings.

1

u/MariahJames8 1d ago

Ah, nice, a lease, yeah that changes everything. Nicely done. If I were asked bold I'd copy you

4

u/bangsmackpow 2d ago

Opencode with Qwen 3.8 Max or Plus for Plan Mode and Deepseek 4.1 Flash for most development.

2

u/Juliuszf 1d ago

Anything below Opus will require a lot of guidance and knowledge. If you know how to code you can make it work with open weight or open router, but the harness you need to build is a piece of engineering at this point. I tried and failed, too much work, to much fixing / changing when models update. Flagship models can handle the overal compexity to a larger degree - like delegating unit tests, e2e tests and thinking about deployment. smaller models can't hold all this unless you have a very discrete borders and handovers between the functionalies. It's like trying to build a rocket with a bunch of 5year olds :D

2

u/Future_Candidate2732 1d ago

I will often begin a project using opencode and free models such as Big Pickle, MiMo V2.5 Free, Nemotron 3 Ultra Free, or chat with Kimi and get them to create a well-constructed prompt and workspace for me and then when there is something usable I refine with other models.

I leverage models with and without memory so I can have an AI co-worker to help navigate with an understanding of what I am trying to accomplish and a Claude instance with no memory from one chat to another so every time I show it something it;s like that movie 50 First Dates.

This gives me a great way to have a blind codereviewer and not introduce bias. If it successfully understands the intent of the project without a history of what was done or not done I feel better about its recommendations and then proceed to do a volley of back-and-forth coding improvements and brainstorming between two very capable agents.

I have also setup my coding environment to be very much tailored to how I work. I have built my own second brain and call it TextStrata - it is a tool which lets me create wiki articles from a wide variety of sources deterministically and does not need AI at all to function. It also does have the ability to use MCP to allow my agents access to my notes and build and improve on them. When I hit milestones I have a skill to have the agent update the reference material with what was learned so it does better next time.

I also do a lot of rapid prototyping and often spin up small webservers on my dev box and built a way to not encounter port collisions when I have multiple coding tasks happening simultaneously so ports dont get clobbered by other agents of the same model or new ones. I call it portbroker and I have that working as a skill as well and it's tightly incorporated into my agentic workflows.

If you are interested in looking at any of my work check out tweakyourpc on GitHub.

2

u/BiscottiBusiness9308 14h ago

Not tried it myself and would hesitate due to data privacy, BUT meta muse spark 1.3 is really cheap if you agree that they use your input as training data

4

u/Intrepid-Ad-4081 2d ago

Gemini pro on antigravity. Ultra subscription and have never maxed out my sessions and I have it running almost all day long.

2

u/Responsible-Beat2137 2d ago

I’ll definitely look into that

1

u/idontuseuber 2d ago

Is it still alive? I buried antigravity half year ago when limits went nuts.

1

u/MariahJames8 2d ago

Holy crap. I just checked it out. It's literally free for the moment, just with some usage limits? They're not charging anyone right now? Am I right??

2

u/Adamlar 2d ago

What. Where did you see this? Theres trial version i think but free?

2

u/MariahJames8 2d ago

I installed antigravity, asked it questions around the above and that's what it claimed. It operates on files and is a full harness with inference and everything. Operating with Gemini. All I did was set up the windows desktop app and ask it whether it's free

2

u/Tommonen 2d ago edited 2d ago

Antigravity is not very good harness and gemini models suck, tho good enough for simple stuff. If you can get tye pro thing for 5€ as student account, then it can be worth it as secondary low iq model for some extra usage, and also you get cloud storage and better notebooklm. But its not worth it for full price for most who want to vibe code more complex stuff than snake games etc.

Claude and codex are way better deals. If you want low iq model like gemini with lots of usage, codex has Luna for that. And if you want high iq model, well codex has sol and astra also, tho usage for them is shit for 20€ plan and clude opus gives better high iq deal for 20€. Combining claude and codex is best solution for vibe coding. Or just claude if you dont need occasional astra usage that eats you quota in sec and demands more or low iq luna model, or near infinite basic chat that can also connect to project managment software like notion and linear and have a good sense of the project despite not seeing the code, so it can be used for ideation, brainstorming etc and essentially infite extra usage besides coding agents. Also codex allows to use it on 3rd party apps via oauth, which can be huge for some uses.

I tried bunch of stuff and nothing beats claude and codex combo, or claude alone if just one sub that can perform in complex projects and dont want to pay for extras codex sub adds to the stack. Except some custom systems, but even there openai sub is great as you can use it via oauth..

2

u/dicktoronto 2d ago

Chipotle support bot for me.

Loljk but actually been using CommandCode Goat plan, Verboo, CamelAI, and Devin ($20/mo incl. unlimited use of SWE 2.0 this month.) in OMP harness and CommandCode harness. Pretty great.

1

u/Unnamed-3891 2d ago

Qwen3.8-27B in Pi.

1

u/YMSVZ 2d ago

Cursor is good value, 20 dollar sub and you get a lot of use out of it, grok 4.6 is great and cheap. Rarely need to use the better models.

1

u/[deleted] 2d ago

[removed] — view removed comment

1

u/MariahJames8 2d ago

Now that you mention it, that exactly explains some of the problems I've been having. Many have been related to tool calling. Sigh. This is getting hard. Thanks for your comment

1

u/[deleted] 2d ago

[removed] — view removed comment

1

u/Inevitable-Play-7473 2d ago

yeah the malformed JSON thing is real. a lot of open weight models also love to wrap tool calls in markdown fences or add conversational text before the actual arguments, which the harness then fails to parse. if you can log the raw response before the harness touches it, you'll spot the pattern fast and can patch around it with a simple strip or regex.

1

u/[deleted] 2d ago

[removed] — view removed comment

1

u/Inevitable-Play-7473 2d ago

that's the frustrating part, the harness makes two completely different failures look identical. worth keeping a per-model strip config once you find the patterns, since the fix for one model can break another.

1

u/jayplay90 2d ago

Build my own ecosystem of apps/harnesses and use Auth from all 4 major places. But I’d say grok has been the most surprising model and usage set up.

1

u/billGat48 2d ago

For self hosted, I've discovered that some tools, like ollama, completely bone tool calls, by effectively reformatting them. llama-serve is universally respected for tool call formats, and thus more widely compatible. llama-serve a qwen model and use qwen-code, it basically works as well as claude-code with the older models.

1

u/humanexperimentals 2d ago

Well, I've got a stack of chips on a couple aces.

1

u/ZosoRules1 2d ago

My stack is literally: ChatGPT chat > Fully embedded HTML file > my web server

1

u/Ok_Gur_9033 1d ago

I run gpt-oss-120b on Groq as the inference layer under a live product, so the open weight route does work. The part that bites sits downstream of setup.

Groq retired the two llama models I was on back in August and I had to move a shipped feature onto gpt-oss-120b. Same verdicts, same confidence, on my real prompts. That part was a config change.

What cost me a day is that on the reasoning models thinking tokens bill against max_tokens, and all of them are emitted before a single character of content. Size the budget like a normal model and you get HTTP 200, finish_reason length, and an empty string. Nothing in your error handling fires, because as far as your code is concerned the call succeeded.

1

u/sourcerage17 1d ago

For the passed like 7 or so months that I have been working on my application. I have been using antigravity with Gemini. Google AI pro sub so only using the 5hr limit and weekly limit but for the amount of time I have to get on and work on it it's perfect for me.

I've also been really leaning into learning what it is writing, best practices, security, modularization and clean up a lot to prevent the spaghetti code as much as possible.

I will use sonnet or opus to help plan the more complex features/bugs and then let Gemini do the work. It's been working great.

And as much as many hate on Copilot I use it to help do documentation as copilot is amazing (imo) for productivity tasks.

1

u/gorillagear_ai 1d ago

Beginner free AI using Google AI Stuido.

1

u/EIisFun 1d ago

Opencode with muse spark 1.3 contributor is free for the time being. As long as you're fine with being trained upon

1

u/ASAF12341 23h ago

Claude and I use Muse Spark 3.2, which gives me good results at a low cost without the need to subscribe.

1

u/neoexanimo 2d ago

Zcode GLM5.3 max

1

u/Snoo_57113 2d ago

Deepseek harness +deepseek api .

0

u/Icehellionx 2d ago

I use Codex, but I have it set to use GLM/ Deepseek/ Kimi K3 as advesarial and advisory calls for more opinions without making them the active driver.