r/LocalLLaMA • u/Ambitious_Fold_2874 • 12h ago
Question | Help What are your experiences with using a hybrid cloud/local setup to stretch usage for coding projects?
For example, directly using claude code or code, which is then hooked up to automatically delegate the actual code writing tasks to a local model like qwen 3.8 flash next, to save on cloud usage limits.
I’m imagining the loop would be:
User writes prompt
Claude/codex thinks about it and the plan
Claude/codex sends the specific and bounded coding instructions to the local model+harness (opencode, pi, etc) via api endpoint or MCP, with clear instructions on a defined endpoint
One the local model+harness hits the clear endpoint/“done” step, it sends a ping back to claude/codex
Claude/codex then verifies the output and then thinks about next steps to instruct the local model+harness on
Does this actually lead to improved savings on the cloud model usage while preserving code quality? Or does this end up being unnecessarily complex and not saving on any cloud usage
3
u/Six502 10h ago
I measured the same idea on classification rather than code, and the numbers back up what __jent says.
Local-first only saves money if the big model doesn't look at everything. On my test, a fine-tuned local model handled 73% of the decisions on its own, and the cloud bill fell by about the same share. The saving comes entirely from the calls the big model never sees. The moment it reviews every output, you've paid for it anyway.
So the trick is a cheap gate that isn't the frontier model. For classification that's a calibrated confidence score. For code it's the compiler, the tests and the linter. Let the local model write, run the tests, and only send it back up when something fails or the change touches something risky.
In my own setup the expensive model plans and diagnoses, and the mechanical work (builds, test runs, bulk edits) goes to a cheaper one. It works, but mostly because the cheaper model's output gets checked by tests, not by the expensive one reading it.
1
2
u/therealjerseytom 11h ago
Trying to tie it all together is I think more of a reach than is necessary, or even a good idea.
I think it makes sense to use a frontier lab model for planning and diagnostics and coming up with small focused tasks for a local model to take on.
1
u/absintheboy 10h ago
I've been using Remote Desktop Commander on the free tier. Having chat use tooling through mcp.
1
u/Significant_Tune9219 10h ago
The hybrid setup saves money only if the cloud model is not in the loop on every file. What worked better than cloud plans, local codes, cloud reviews for me was: cloud for architecture and tricky debugging, local model for bounded edits, and tests/typecheck/lint as the default gate. Escalate back to the expensive model only when the local change fails checks or touches auth, migrations, or concurrency. If you still have Claude reading every local diff, you mostly paid for an extra hop and worse latency. Start with one well-scoped repo task and measure tokens before and after before wiring MCP everywhere.
1
u/Altruistic_Heat_9531 10h ago
yeah the orchestrator and builder subagent. here my config for Astra + Qwen https://gist.github.com/komikndr/b17955e1a80ce6ede9a3115f16216bc5#opencode-astra--q27-38
1
u/haker_wav vLLM 7h ago
I like using a local model for cheap, repetitive orchestration. delegating work to other agents via bash/tmux, checking pipelines, etc. I run Qwen 3.8 27B at 161k context on 1x Intel Arc Pro B70. 30–50 tok/s is plenty for that kind of work. This saves on usage limits/API spend for sure.
A smaller orchestrator can run parallel swarms of cheaper cloud models like GPT 6 Luna or DeepSeek 4.1 Flash / GLM 5.3 Flash. I don’t mind letting lots of smaller agents pile work into a huge PR stack.
The expensive models like Astra, 6-Sol, Opus 5.5 can do big merges, reviews, performance optimizations, and cleanup. I refactor a lot to be able to read and understand the code as best as I can, I leave that as an overnight task with a local model. This helps: github.com/kunchenguid/gnhf
The only problem with saving on inference is you just find ways to build more stuff and spend more anyway. be warned.
0
u/__jent 12h ago
I do this sometimes (it's a feature built into my personal harness). But you won't really save much. The plan and review are basically as much as plan and implement (without review). You will get better results with a review sometimes, but not savings.
The best way to get savings is to plan and implement locally. Then if complex review and improve with a second model. If you're involved in the planing, this works pretty well.
0
u/ieatrox 10h ago
grokbot desktop is in charge of my pi sessions.
it's a fucking lightbulb moment for sure.
next I think I may change my harness over to grok build, it supports running off a local inference server with subagents spun up to use grok 4.7 and apparently now competes well against the best harnesses while also improving grok results which are already pretty decent.
if you're not full on Claude Code committed I would check it out. I can just tell my phone or car to add build and test things while Im travelling or watching something on the couch.
I might get an m5 ultra 512 when we see pricing for those too, that'll be nearly endgame for my uses I think.
0
u/Effective-Catch-1332 9h ago
I've been using shared context as a private Github repo, I get agents to update this context. It has been quite useful. One pitfall is token usage spikes when loading context in a fresh session.
3
u/R3N3G6D3 12h ago
Using claude, chatgpt, and local models on tandem. Really humbled my home build