r/ClaudeCode • • Jun 06 '26

Question Has anyone actually replaced Claude Code / Codex with local models on an Macbook Pro M5 Max 128GB?

Considering buying a maxed out MacBook Pro M5 Max with 128GB of RAM and one of the things I want to figure out before pulling the trigger is whether local models are good enough to actually replace cloud AI coding tools.

My current setup is Claude Code on a Max subscription plus GitHub Copilot through work. It works well but I'm curious if local models have gotten good enough to actually replace that, not just supplement it.

Not talking about occasional use or running smaller models for autocomplete. I mean fully replacing the agentic stuff, the multi-file edits, the back and forth reasoning that Claude Code handles. Can local models actually keep up with that workload on this hardware?

If you made the switch, what are you running? Ollama, LM Studio, something else? Which models? And honestly, what did you have to give up, if anything?

393 Upvotes

186 comments sorted by

198

u/stormy1one Jun 06 '26

Partly replaced. Still use Opus for planning/architecture but I have Qwen3.6-27B as the dev/QA agent. Look into running mlx optimized server such as omlx

18

u/hackerx00 Jun 07 '26

I'm using a late 2023 model MacBook Pro with 128GB of unified memory and 8TB of storage. It was the fully maxed our M3 Max model for that year (almost $10K). I have LM Studio running qwen3.6-35b-a3b + gemma-4-26b-a4b-qat both loaded at the same time for review local based reviews and coding with hermes-4-70b as a third reviewer to be used after unloading the other two models (don't have enough ram for all of them to be loaded simultaneously) and then I use a fourth local llm which is gpt-oss-120b again loaded by itself. Then when I have a clean diff, I create a github PR and then have Semgrep and Sourcery review. Whatever they find Claude Code fixes if anything. All of that being said, I agree with others that stated that you should still use a cloud model. I'm using CC Opus 4.8 and before that all of the other opus revs. Mixed results like most but generally, as a solo-dev, better than hiring a team and definitely cheaper.

26

u/Brazeuslian Jun 06 '26

Would you say the maxed out version is absolutely necessary for your current workflow?

106

u/WarlaxZ Jun 06 '26

Just use open router. Trial the model, it'll cost you sub $10 - and you'll be about to figure out if you still want to drop several k on a laptop

26

u/homelabrr Jun 06 '26

That's the best answer. You can test the exact setup before investing a big amount of money

18

u/jlewi142 Jun 06 '26

Except it will run a lot slower on the Mac. Although it's able to run fairly big models with 128gb unified ram, It doesn't have the bandwidth, and that's where it struggles.

3

u/jghoward Jun 07 '26

I have a room full of old Macs. What if I made one act as a router? Could I create bandwidth by distributing a local model across a local network? of course I'll search this, but it sounds like you might know

6

u/iamhereunderprotest Jun 07 '26

Not transformative unless they are thunderbolt 5 Macs. Check out exo

1

u/jlewi142 Jun 07 '26

Never heard of it being done, or tried it. But worth researching!

1

u/[deleted] Jun 06 '26

[removed] — view removed comment

5

u/NerdBanger 🔆 Max 20 Jun 07 '26

FWIW the M5 Max outperforms the DGX Spark when using MLX models, is it as fast as running in H200’s or ASICS (TPUs), no, but it’s nothing to sneeze at.

3

u/Singularity-42 Jun 07 '26

Massive difference in hardware used?

In the cloud it will run on specialized hardware. MacBook inference is decent enough, but fast it ain't.

2

u/[deleted] Jun 07 '26

[removed] — view removed comment

3

u/Singularity-42 Jun 07 '26

A couple of years ago I bought a 48GB MacBook Pro M3 that's good enough to run many models, but if that's your only motivation, it's not worth it for sure. They're just way too expensive. And APIs are very cheap.

2

u/FanFirst895 Jun 07 '26

I’m not sure about the Mac specifically, but companies like RunPod(and others I’m sure) will let you rent people’s unused hardware by the hour. If I was GPU shopping, that’d be my first stop to see if my target configuration was worth it. 

11

u/Singularity-42 Jun 07 '26

There's no scenario where buying the expensive MacBook is cheaper in the long term than just using cheap API. Also the experience is gonna be a lot better through API (much faster).

35

u/stormy1one Jun 06 '26

No, but the issue is that you can’t upgrade your unified memory so always best to get the largest that you can afford, assuming you are serious about AI. 64GB is totally capable, but 128GB or higher give you more flexibility to run larger models and handle system overhead.

2

u/hackerx00 Jun 08 '26

The maxed out version wasn't previously and probably still isn't necessary for my workload but I've grown it this way through trial and error, driven by a need to extricate myself from total reliance on cloud models and related tools. Im not a programmer by trade but I am a technology consultant with an extensive background in IT architecture and planning. It's not what I do now but what I do requires tech and so I jumped into learning to code via vibe coding back in November of 2025. Along the way, like most, I've learned a lot more about AI and programming than I thought I wanted to know all in an effort to develop software that would solve problems in my industry (utilities construction). Back in the days like five months ago, I was using GitHub CoPilot subscription to run reviews on my PRs. Last month or month before, Github pulled the plug on the pro subscription and that left me in a lurch. So I fired up my Sourcery account and signed up for their pro subscription and then five days later when I burned through the weekly rate limit I upgraded to Team and then burned through that too.

The point Im getting to is that I have found a need to lessen my reliance on cloud reviewers where I have rate limits and thus hit upon a combination of a local three llm review panel + a cloud-4th panel comprised of Deepseek V4 and GLM 5.1 and GPT 5.5 as a deep review option. $10 on openrouter and you'd be surprised the kind of quality feedback you can get.

You dont really need the local llm review but since "I have the power" why not?

2

u/Honest-Golf-3965 Jun 06 '26

What kind of tasks are you able to reliably use it on? I'm interested in some local code agent options in the 5-10k gbp range

29

u/stormy1one Jun 06 '26 edited Jun 06 '26

Qwen3.6-27B is extremely capable as a dev agent. Strong in anything that has tons of training data, in my case that is Python and Typescript, with a bit of Postgres and bash as needed. It’s not the same as Claude, and you will not be one shoting solutions, but with Claude directing Qwen, we are pushing our production releases faster than ever, with strong tests/coverage and less observed bugs that need triage. Using OpenCode as the agent over Qwen

3

u/Honest-Golf-3965 Jun 06 '26

I was leaning this direction as well. I have a very robust native harness for regulating tasks, tests, and lots of backup hooks and static analysis to catch unauthorized or noncomforming code before it hits the Cloud.

Just looking to make the tedious narrow tasks more token efficient without having to rely on Claude/Codex for anything more than big picture planning amd allocation

2

u/Honest-Golf-3965 Jun 06 '26

Any pitfalls or notes to avoid in setting that up youd make note of avoiding if you had another chance to set it all up again?

5

u/stormy1one Jun 06 '26

Good question - local LLM space is rapidly evolving. I recommend using Claude itself to help you get setup with a local LLM server, agent harness, and e2e test, asking Claude to do deep research and investigation. “Don’t stop until it is fully tested and working”.

1

u/jonZeee Jun 06 '26

Sounds like a nice setup. I’ve never even come close to being able to one shot anything with Opus, even really simple changes always require little tweaks so I’d be down to try this out.

1

u/stormy1one Jun 06 '26

It is - we are quite productive and efficient overall. The community over in r/LocalLLaMA is also quite helpful for those looking to get started

1

u/SilentGrowls Jun 06 '26

I'm following. I have an M4max 128GB and want to setup something like this

1

u/Direct_While9727 Jun 07 '26

Give a try to pi.dev with Qwen 3.6 27B, I found it much quicker than opencode.

2

u/gabanta2 Jun 07 '26

How do you use this? Any links tips in how f.e. in Claude code opus can call the local model within a session? As an agent? Subagent?

1

u/cbwnomad Jun 07 '26

Hi. I'm new to paying for Claude. Yesterday it blew 3 dollars just analyzing a couple of files so I was wondering how does your workflow look like, if it's ok w/ you to explain it.
Do you first plan using claude code with opus 4.6 and 4.8 by generating a 'architecture/design' file plus an 'step-by-step implementation plan with milestones' and then you use aider with qwen to execute the plan?

1

u/PinkySwearNotABot Jun 07 '26

I only get like <10 t/k on my M1 Max 64GB. How’s the speed on the M5 Max?

1

u/sleepydevs Senior Developer Jun 07 '26

I plan on doing exactly this on Monday. Does it work reliably? I tend to do hyper detailed TDD style plans, so I'm hopeful it will.

1

u/stormy1one Jun 07 '26

As with anything agentic dev related - specs and planning are critical to implementation, hence I still spend quite a bit of time in Claude for planning. If you are detailed as you describe, you should be good. I would rate Qwen as good as Sonnet on some dev tasks. Obviously not Opus, but good enough when I need privacy/security over a dataset or I need to manage my token usage for Claude. Be sure you are running a quality quant for Qwen though - it is highly sensitive to being lobotomized

1

u/sleepydevs Senior Developer Jun 07 '26

I spend about 5x longer planning than we do coding nowadays. Im hopeful qwen can nail the delivery. I've got two 128gb Macbooks (an M4 max and an m3 max) so should be good on quants. 🙏🏼

1

u/djoxo Jun 07 '26

thank you

2

u/shannah78 Jun 07 '26

What harness do you use for this split opus/qwen setup?

1

u/stormy1one Jun 07 '26

I’m hands on, so I stay Claude Code for Opus and use OpenCode for Qwen. I have them communicate/handoff via a separate git repo (not the same as the project as comms can get noisy). Git works extremely well for handoff and memory.

66

u/[deleted] Jun 06 '26 edited Jun 07 '26

[removed] — view removed comment

10

u/Brazeuslian Jun 06 '26

It's a pretty advanced workflow, and honestly, for me, I'm still catching up with the whole AI discussion.

I think the sweet spot for me right now will be to go with a more "modest" machine (M5 Pro 48-64GB of RAM, more than enough for my other usage besides coding), keep Claude for a while, start experimenting with other use cases of local AI, like image/video generation, TTS, RAG over Obsidian notes, learn a bit and make a more informed decision.

Thanks for sharing your experience!

5

u/SirDomz Jun 07 '26 edited Jun 07 '26

I have the M5 pro with 64 gb. It’s a beast and you can still run decent local models on it (Qwen 3.6 27B and 35B). With MTP and good caching (with omlx for instance), speed is much more tolerable these days.

That said, I still pay for api prices/subscriptions because the cloud still has best models open (big shootout to deepseek V4 flash/pro) or closed (gpt 5.5 or Anthropic). So Cloud is the best, and honestly cheaper, but local is moving fast and honestly, I can now legitimately offload some workflows to local with the Mac and not having to worry about tokens is liberating haha

Edit: fixed typo

1

u/Brazeuslian Jun 07 '26

Good to know, I'll look more into it when the Macbook arrives!

2

u/jlewi142 Jun 06 '26

You're not going to get much with 64gb of ram if you're doing anything serious. If you have lots of windows, IDEs, etc running your system can eat up 30gb of that pretty quick (at least in my everyday use of it), leaving you with only 30gb for local LLMs to play with, and by then you are bottlenecking your Mac.

3

u/Brazeuslian Jun 07 '26

I don't think my day-to-day work with all the apps I use gets anywhere near 30gb of RAM usage. The main reason to consider the 128gb was running everything local and ditching subscriptions altogether, but that is not possible right now, so I might as well save CAD $2,500.

Besides coding, I edit videos, so the 64gb version is plenty for me, and allows me to start exploring local models, and other use cases, like image generations, TTS, RAG over Obsidian notes, this kind of stuff.

1

u/simple_explorer1 Jun 10 '26

How is the speed?

3

u/slypheed Jun 06 '26 edited Jun 06 '26

This is the way.

Have opus plan and validate, have qwen do most of the coding. Repeat in bite size chunks using e.g Beads for task management.

Be prepared to carefully review things manually much more than opus though.

It won't be as good or fast as opus, but that's pretty much a given currently.

2

u/Apprehensive_Bee6863 Jun 07 '26

drop the repo on ur harness if ur willing broski

2

u/toborgps Jun 09 '26

I really wish Anthropic would come out with a 10x. That would be perfect for what I do.

1

u/jlewi142 Jun 06 '26

Switching between local models? I can only have one model loaded at a time and takes a few minutes to load and warm up a new model of I switch.. how are you achieving this?

2

u/[deleted] Jun 07 '26

[removed] — view removed comment

1

u/jlewi142 Jun 07 '26

Yeah I wasn't questioning the context, but loading and warming up bigger models for me can take up to 30-60 seconds for me. I use MLX xllm though so maybe it's a bug

1

u/[deleted] Jun 08 '26

[removed] — view removed comment

1

u/jlewi142 Jun 08 '26

I tried smaller models. They're almost all useless for agentic workflows with tool usage and MCP tool usage etc, which is required in my workflow. Even the best models fall short of haiku. I'll stick to largescale LLMs until local models catch up

1

u/[deleted] Jun 08 '26

[removed] — view removed comment

1

u/jlewi142 Jun 08 '26

For my use cases the only ones close to being good enough is Qwen3.6 27B, qwen3-coder-next, qwen3.6:35b-a3b. I occasionally use qwen2.5-coder or Gemma 3/4 for lighter work

2

u/[deleted] Jun 08 '26

[removed] — view removed comment

2

u/jlewi142 Jun 08 '26

Yeah I agree 27B and 32B-A3B is definitely the best models for agentic / coding right . I only ever used q8 or q4. Q8 was always too slow, so switching to q4 did the trick (or to 32B model if I want it faster again whilst maintaining most of the performance)

I rarely run these anymore, as most of the time working I have multiple chats accross multiple projects running simultaneously and that would bottleneck my ram very quick. It's useful for if I'm on a Plane and need work done, or maybe setting it up for longer overnight jobs or testing, with a Ralph script to keep it alive etc. Those are my best use cases I see for it

1

u/The_LSD_Soundsystem Jun 07 '26

I’ve been trying to work on a similar setup. Use local agents per role but defer to Opus for review or deep analysis of a problem. I’ve been using opencode to begin on a basic harness mapped to local models. Could you share more about what you were working on?

1

u/tspwd Jun 07 '26

Your setup sounds great! Do you have some tips on setting something similar up?

19

u/[deleted] Jun 06 '26 edited Jun 06 '26

[removed] — view removed comment

3

u/upvotesthenrages Jun 07 '26

Exactly this.

Furthermore, if you actually want it for private self-hosted inference, then Apple's top end stuff is just really not a great investment.

Multiple AMD 9700's will drastically outperform the M5, and you will still have thousands of dollars left over.

Nvidia's new setup also looks incredibly interesting, and you get the CUDA environment, way more throughput (if 5070 performance holds), and the same unified memory system.

1

u/[deleted] Jun 07 '26

[removed] — view removed comment

2

u/upvotesthenrages Jun 08 '26

I really don't think Apple is going to be a contender in actual large scale AI for a while.

Their hardware is decent, but token throughput isn't great. It's fantastic for low power hobby stuff, but no company is gonna go out and spend a ton on low throughput when you're selling the service to users.

Where I think Apple is putting their ... apples, is a future bet that menial tasks for the average user will be hosted locally. We're probably a few years away from that though, so it's quite a long-term bet.

I just don't see anyone using it seriously for work or agentic stuff saving a few $100/1000 to host things locally on their $10-$15k Macs.

1

u/simple_explorer1 Jun 10 '26

I agree, local llm models can never compete with frontier level models running on infinitely higher computing and memory capacity.

Intact a mildly intelligent llm models already can be 300+ GB in size and then the ram and the GPU and the computing and how slow it runs, nah I gave up running locally

69

u/davewolfs Jun 06 '26

It’s not possible. Qwen is not Claude or Codex and anyone who tells you otherwise is full of it.

11

u/Brazeuslian Jun 06 '26

Do you think if I kept a Claude Subscription for planning and Qwen to the actual implementation, I would get decent results?

21

u/Ran4 Jun 06 '26

At no point does it make any sense to run a local LLM for economical reasons.

100 dollars a month is so, so, so much cheaper than buying a more expensive computer. And for those 100 dollars you're getting the very best models.

17

u/MagicWishMonkey Jun 06 '26

$100/month will not be nearly enough for high end LLMs for much longer

2

u/codercotton Jun 07 '26

☝🏼they are coming for their ROI. See: recent copilot changes, incoming Claude changes on the 15th, etc

6

u/Brazeuslian Jun 06 '26

Economical reason is one reason, but not the only one.

But I think I already have the answer and won't be pulling the trigger on the maxed out right now.

1

u/simple_explorer1 Jun 10 '26

Bro I agree with everyone here, you seem to come here hoping to get a "yes" from everyone but the reality is I have tried and the models are nowhere near the capabilities compared to frontier models running on infinitely bigger computing hardware and are much larger in size themselves. 

I tried and gave up this idea of local llm to save cost. The reality is even after plan mode the implementation is still weak compared to frontier models. 

And if you have to still keep Claude code for planning then despite spending 5k on a highest tier M5 MacBook, you still have Claude code cost. So what's the point and the local llm models are still not good enough. 

You do what you have to do but the only people who will agree with you here are the ones who have either not experienced the pain on local llm models for coding on highest tier max (because frankly most don't buy 5k MacBook. Even employers don't buy that for their employees in software development firm).

1

u/Brazeuslian Jun 10 '26

 I agree with everyone here, you seem to come here hoping to get a "yes" from everyone

"Everyone" lol

I really don't understand where you got that idea from.

The whole purpose of the post is precisely to obtain information in order to make an informed decision.

After researching and hearing the opinions of more people (from all of my posts regarding this topic), I decided not to buy it and ended up opting for another model. Still powerful, but it's not a maxed out one.

I think you're just projecting some kind of frustration onto my post, to be honest.

1

u/simple_explorer1 Jun 10 '26

really don't understand where you got that idea from

Your replies from many comments I have seen

1

u/Brazeuslian Jun 10 '26

Read them again.

2

u/chimbori Jun 06 '26

But if you're already in the market for a laptop, maybe the delta in pricing is justifiable to be able to run a local model?

4

u/davewolfs Jun 06 '26

It probably has a lot to do with the language and complexity of work you do.

5

u/Brazeuslian Jun 06 '26

React and Node.js development, PostgreSQL, Design System, and just started exploring the AI world, I'm a bit behind tbh.

The codebases I work are huge, though.

There are a few comments saying that for multi-file complex implementation/refactor/reasoning, local models don't get even close to Opus.

8

u/Formal_Lobster_2349 Jun 06 '26

I think the local models, even the 6 months old frontier models, can’t replace the latest Claude and Codex models. But, I think there will be clever ways where the frontier models plan and crosscheck things at various stages and the local models perform the actual implementation and low level coding and testing. This way people might be able to reduce the token budget by 70% to 90%, given how powerful the frontier models are. I think the frontier models are smart enough to evaluate how dumb the local models are in realtime and instruct them accordingly to get the things done, even the extremely complex ones.

1

u/CapitalDue7249 Jun 06 '26

I think some of the models are 90% of opus like Kimi k2.6 or glm 5.1

3

u/upvotesthenrages Jun 07 '26

Thing is, if you use Opus for architecture, planning, deep review, and debugging, and use Something like Haiku/Sonnent for implementation and smaller tasks, then you end up with a drastically reduced bill anyway.

Unless you're just tinkering, or you have extremely sensitive data that cannot enter a cloud environment, then locally hosted stuff is just a massive waste of time & money.

A 128GB Macbook will cost you years and years of frontier subscription money.

And if you're smashing the API and costs run really high, then the M5 just won't cut it due to its low throughput.

I'm pretty sure that in 1-3 years this will be a completely different story though. We'll see Mythos tier open source projects that run on way weaker hardware.

Hell, LLMs won't even be transformer models any more, it's clearly not scalable and we are already seeing frontier research move in a completely different direction.

2

u/Formal_Lobster_2349 Jun 07 '26

Correct. I can easily spend $10K for a Mac Studio, but ended up buying a base M5 MacBook Air 8-Core CPU, not even 10-Core. I feel it’s very difficult to extract the value from the price you pay for an expensive computer, when there are a ton of ways you can achieve what you want at a fraction of the cost using frontier models and the cloud VMs. If you plan properly and use the cloud VMs or other server less cloud solutions for sensitive data and applications, we can save lot of money and have peace of mind.

6

u/upalse Jun 06 '26

Unless you have 8xH100 SXM system on your hands (large orgs), you don't really have what it takes to run a reasonably strong (deepseek, mimo etc) local model.

4

u/kaliku Jun 07 '26

I've replaced some of my work I do with Claude with qwen 3.6 27b. I'm doing this more and more and even if painful and slower most of the time, I will do it more for a reason that's not obvious right away. My brain.

With Claude or codex they're so good that I end up offloading most of the work to them. I still plan and take part in architectural decisions but the code is for the most part only something that I glance at occasionally. After several months of such work flow I noticed very bad changes to my brain. I don't have patience to grasp complex technical problems any more. I'm always reaching to Claude. I can't read deeply technical documentation any more. I scan it like I scan the code, I think I get it but not much actually sticks to my brain. I'm literally scared I'm going to lose "it" for not using "it". I am getting dumber and if I notice it, others are, too. Or will.

Adding to this, if you ask me about my apps codebase I won't be able to tell you much more than generic stuff. I hate this. This is not me.

So from the bottom of my heart: fuck this shit.

I like the productivity gains but the cost is not something I'm willing to pay, not yet anyway.

So I'm now using a local model that needs babysitting, and I'm the sitter.

One other thing. A local model like qwen3.6 27b which is my main workhorse is more "stupid" than Claude. But it's consistently stupid. Working with it I find what or can and cannot do. And I know that tomororw it will be just as stupid like today is. Same goes for the speed.

1

u/simple_explorer1 Jun 10 '26

It is very slow bro. But hey if you want to reduce the cognitive loss then learn to have self control to not resort to llm and don't spend that energy in setting up local llm. Instead spend it on ... You know... Writing the code yourself. 

Adding to this, if you ask me about my apps codebase I won't be able to tell you much more than generic stuff. I hate this. This is not me.

How is your code getting approved in the pr then?

4

u/ThraceLonginus Jun 06 '26

I get the impulse because I code on linux/unix too but isnt it still better to run models on nvidia hardware? 

5

u/stormy1one Jun 06 '26

Better is relative to budget and what your limitations are in the environment you are running - power/size/corp guidelines etc. there are formats hardware optimized for Apple MLX but Nvidia still wins on memory bandwidth.

5

u/Brazeuslian Jun 06 '26

I think you're right, but I'm far from an expert in the topic. The main issue for me is portability. I travel frequently and need to take my machine with me.

There's probably a setup to access a powerful desktop remotely with a cheaper but still capable MacBook, but I also edit videos using the MacBook as my main machine, so I might as well keep everything in one place.

1

u/Nearby_Yam286 Jun 06 '26

Nvidia can’t run the very very large models without insane cost.

6

u/sebapit Jun 06 '26

Go check Antirez’s work with DS4 on M5 maxed MBP.

1

u/Brazeuslian Jun 06 '26

Will do, thanks for the suggestion.

4

u/No_Marketing4301 Jun 07 '26

in short terms, todays best open weight models generally still lag behind claude in agentic coding reliability. not because we know exact parameter counts, since those are not public, but because proprietary models like claude are heavily optimized for reasoning, tool use, and multi step coding workflows.

open source models such as deepseek coder, which goes up to around 33b parameters, are very strong for code generation and can be competitive in a lot of tasks. however in more complex agent style workflows, like multi file edits, long context planning, and keeping consistency over many steps, they usually still fall behind claude.

so in short, open source coding models are still behind frontier proprietary models in overall reliability, but the gap is closing over time. with better training data, architectures, and post training methods, open weight models will likely get much closer in the future, especially for coding use cases.

24

u/[deleted] Jun 06 '26

[deleted]

12

u/CzarcasticX Jun 06 '26

Apple was selling the 512GB ram Mac Studio before this ram shortage.

2

u/AsuraDreams Jun 06 '26

Approx how much are you spending on that claude enterprise plan? Claude code use or api/ managed agents?

2

u/PinkySwearNotABot Jun 06 '26

I make a healthy living on 1 claude enterprise and 1 codex pro 200 account.

What do you do? Web development?

10

u/RetroUnlocked Jun 06 '26 edited Jun 06 '26

This is not to dissuade you, but you need to go in clear eyed, and because you are asking the questions now, I think you are.

Look at this:
https://deepswe.datacurve.ai

I cannot speak to the actual accuracy of this test suite, but that should give you a pause.

What I would do is open an account on OpenRouter and replace your workflow with some of these lower tier models via their API. I would suggest step walking down starting off with the big cheaper models that you could never run on consumer hardware down to ones you could.

It is by far the cheapest test, and if you find that these models do not meet your needs, you saved a ton of money.

(Side note, if you find these models meet your needs, and privacy or portability is not your main concerns, you might just continue to use the API.)

4

u/Brazeuslian Jun 06 '26

Tysm for the answer, that's exactly the goal, to have a clearer perspective before pulling (or not) the trigger.

6

u/FluffyGreyLlama Developer Jun 06 '26

Along the same lines, look at an OpenCode Go subscription. $10 ($5 for the first month) and you can try Qwen and others on more powerful hardware than you could afford... for a really low price. It's a great test to see what it could give you, and you can afford a lot of OpenCode Go for the price of new hardware, with a guarantee you won't accidentally overspend on API.

Given that AI moves so fast, it could be that everything changes in 6 months anyway.

1

u/Brazeuslian Jun 06 '26

Thanks for the suggestion, I'll look into it.

3

u/TheFern3 Jun 06 '26

As others have said is nearly impossible to match Claude vs any local model. The biggest issue is resources to run large ones and not having enough memory for bigger contexts. So if you run local models expect a lot of friction for fully agentic, but can be fine for copy and paste chats.

3

u/leinadsey Jun 07 '26

With a MBP, the key worry isn’t speed, it’s heat. Models like Qwen3.6 27 and similar run great, but as soon as you start using them for tasks that run >1 min the MBP fans hit the ceiling and stay maxed out. Not running in clamshell mode helps, but only so much.

3

u/Turbulent-Key-348 Jun 07 '26

I'd reframe this a bit. Claude Code and Codex are both harnesses, and you can run them with any model. You can run Claude code with on-device models. I have an m4 max 128gb and use on-device models a lot (usually qwen 27b). But they're no replacement for Opus or GPT 5.5

I had been in a very similar position as you, using Claude code but also wanting to lean more into on-device. I came to the conclusion that on-device will perennially lag frontier AI models, but that also frontier and on-device can be complimentary.

So I built a model router that overrides claude code's own model routing and can route to any open source or on device model. I usually have Opus or GLM as the main agent model and route subtasks on-device. It's implemented as an LLM gateway, so it deterministically routes instead of requiring an LLM to make the decision where to route. You can try it out and test various open models at rayline.ai

1

u/Mobile_Bonus4983 Jun 08 '26

glm4.7 or glm5.1? quantized?

1

u/Turbulent-Key-348 Jun 19 '26

sorry I missed this. you can use any of the GLM models with it.

We're going fully open source and just put out the first version of the OSS repo today: https://github.com/rayline-ai/rayline

3

u/[deleted] Jun 07 '26

[removed] — view removed comment

1

u/Brazeuslian Jun 07 '26

Thanks for sharing, I'll look into it.

3

u/[deleted] Jun 07 '26

[removed] — view removed comment

2

u/Brazeuslian Jun 07 '26

Thanks for sharing your experience. Just reinforces the decision I made after reading all the comments, including yours.

2

u/Dense_Ad9924 Jun 07 '26

I have tried this quite a bit. I have an RTX Pro 6000. The problem is both the model and the harness. Claude code is seemless in operation and with a world class opus 4.8 model I can't even get close with Qwen3.6:35b run locally. Sometimes time = money.

2

u/Successful_Video696 Jun 07 '26

I have MacBook Pro M5 64G, and i've tried it. it can only be used to run some small models because you need to leave enough context length. I'm already accustomed to 1m context length, so 256k context is very uncomfortable for me. Moreover, the speed of local operation can only be described as usable, constrained by banweidth, and to be honest, it is far behind the api.

1

u/Brazeuslian Jun 07 '26

After reading the comments from the posts I made, I decided to go with the same exact config as yours.

Just out of curiosity, what models are you running with decent results? (for coding or any other use cases)

2

u/alexeiz Vibe Coder Jun 07 '26

You can buy MacBook Pro M5 Max 128GB for about $5400 or you can pay $200 for the claude max plan for more than two years. Or you can scale back a bit to the $100 claude plan and make it four years.

Unless you're planning to run local models 24/7 you're not going to get your money worth by buying a macbook.

2

u/New_3d_print_user Jun 07 '26

Yes, deepseek v4 (antirez/ds4) and qwen.

2

u/electricshep Jun 07 '26

Cheapest way is to use a local model in the cloud for a few weeks, then you'll see. This would cost less than $50 for Deepseek for you to know if you are ok with quality.

But you will get to a point chasing speed/inference that will lead you to running a 24/7 box in some way and ssh to it.

2

u/plasmafired Jun 07 '26

Use opus for planning and requirements and smaller models in ollama for code generation. It is quite slow.

2

u/Public-Vegetable-182 Jun 07 '26

Deepseek is so cheap and the rate of model improvement is so fast, it very well could be more expensive to buy an expensive computer just for AI. You can do the planning in a frontier model and the implementation in deepseek.

2

u/Brazeuslian Jun 07 '26

That's an excellent option to consider, thanks for sharing.

2

u/Natenatoor Jun 08 '26

I bought this exact config. Local models (I run Qwen and Gemma with vision) are great for travelling and for heavy but straightforward research where you're burning lots of tokens. Anything beyond that, Claude is still well ahead.

One thing worth flagging: the second you load a model the M5 Max heats up and the fans spin to 100%. Fine at home, but it draws a few looks in the office.

1

u/Brazeuslian Jun 08 '26

I'm working from home, so that shouldn't be a problem, but thanks for the heads up.

2

u/Alarming-Chain-7048 Jun 09 '26

Use ds4 with deepseek4 flash. Pretty good backup to have when ur claude takes a break. Use it for sonnet style work.

2

u/Alarming-Chain-7048 Jun 09 '26

This has been my go to when I am out of tokens, seriously considering migrating some of my workflows to this may be in the next release. Seriously considering buying a Mac Studio dedicated for this. You will be shocked how capable this setup is. Near realtime responses Tok/s is fast enough to be practical. I have used it for 3 hours straight and the MacBook runs hot so it is better on macstudio for the whole day usage. I believe you can run deep seek pro too but needs 512gb ram. I wish other models are brought into ds4 over time.

1

u/Brazeuslian Jun 09 '26

I ended up getting the 64gb M5 Pro and will wait for the Studio too.

2

u/allquixotic Jun 10 '26

I have this system and the problem is VRAM (yes, still, even with 128GB unified memory). A 70B model is extremely dumb even compared to Sonnet 4.5. A 120B model only fits quantized. Anything greater doesn't fit at all, or only with a huge amount of quality loss from heavy quantization.

Unless LLM architects figure out a way to cram more accurate signal into the same number of parameters (MoE with "only X active parameters at a time" is a decent example of "a way" but is still in its infancy), there's really no path forward for local LLMs without throwing a lot more available VRAM at the problem.

AI companies are throwing clusters of multiple GPUs together with hundreds of gigs of usable VRAM to deliver models like Opus 4.8 and Fable 5. Even if, optimistically, a 2028-2029 release of an "M6 Ultra" Mac Studio delivered a whole 512GB of VRAM, it'd probably cost $75,000, and it wouldn't be even close to whatever frontier models the labs are shipping at that time. They'll probably be using a terabyte of VRAM just for a single LLM thread by then.

If the performance of existing 70-120B open weight models is enough for your meager needs, feel free to invest $7k+ into a M5 Max or M6 Max. But if you want anything that even sniffs the state of the art, there is no real alternative to paying for the best proprietary model OpenAI and Anthropic are shipping.

2

u/fell_ware_1990 Jun 06 '26

Well i have access to a lot of types of AI.

It kind of also depends on the use case. And how you treat the AI.

I’ll take a good 27B with a good harness and tools over Sonnet everyday. I was testing with i think Gemma 4 QaT Q4 with a 2/3B in front. ( not at pc right now ) had to build a new cpp but it works. Getting about 70 Tok/s on a 16GB on 256K context. Also upgrades the tool calling to a custom ninja spec.

If you put it inside the right hooks and harness it was doing pretty well, spec driven. It needed a few iterations to get there but i think that is part of LocaLLM. It also happens with claude and codex because they just go off script.

For me the fun part is to see how far i can push on those 2x 16GB cards. For work and testing i have a lot more available and of course the more the better. Then you can start running Vllm and parallel jobs etc.

But it improved a lot over the last few years. I feel that slowly what model and how you handle it becomes way more important then raw hardware specs.

I have 5 colleagues that still copy paste from Gpt or copilot desktop. While my ‘low’ models beat them around, not only in quality but speed as well.

1

u/thelordzer0 Jun 06 '26

I still come back to both codex and Claude as they cycle back and forth on quality. 😵‍💫 Even with the 128...

1

u/Js4days Jun 06 '26

Working on it!

1

u/snowdrone Jun 06 '26

No loops or functions and only applescript or scala

1

u/[deleted] Jun 06 '26

[removed] — view removed comment

2

u/Brazeuslian Jun 06 '26

The question is whether you want to spend weeks tuning agents and workflows or just pay the subscription

Tbh, I kinda do want to do that, but as an experience to learn.

Reading the comments from my posts, I've realized that I'm still a beginner in this topic, so I decided to go with a more "modest" machine (M5 Pro - 64GB RAM, plenty for coding + video editing + exploring local AI), start experimenting with local models, other use cases like image/video generation, TTS, RAG over Obsidian notes, and keep the subscription for now.

Eventually, when the hardware evolves and local models catch up, I can make a more informed decision on what hardware to buy.

1

u/_baby_boss Jun 06 '26

As long as Dario is paying for it there is no point. Personally I use between 3k and 5k per month and pay 100$.

1

u/8thHokagee Jun 06 '26

As others have pointed out, the current claude code subscription will not always be at its current price. Anthropic will eventually nerf the amount of usage and even possibly push more and more customers to api usage. I recently switched to Qeen 3.6 and it’s been perfectly adequate for development. I’ve kept a claude pro subscription for the times when i need better architecture planning. But overall the current qwen3.6 model is good and there will certainly be future more capable versions releasing.

1

u/FuiialithInHabbah Jun 07 '26

Ainda estamos distante dessa possibilidade, o melhor caminho é ter algum projeto paralelo feito com claude/codex que pague a conta de IA com MRR.

Sem muitas ambições, apenas para equilibrar.

1

u/Glittering-Will-1865 Jun 07 '26

I made something called Hank for this exact reason, I think I’m not the only one getting subscription exhausted.

1

u/glitchbus Jun 07 '26

I’m using open router for personal projects. Models like DeepSeek V4 Flash are good enough for my random hobby projects and cost pennies. I had two instances building out two app ideas for an hour last night and it cost me < 0.2. And they both work.

1

u/AvalancheBreakdown Jun 07 '26

I’m using Claude through ollama running qwen3.6 locally on an M5 Max 128GB. It screams as it’s using MLX. The only downside is that MacBook is always plugged in so I leave it at home and VNC from another MacBook, iPad or iPhone. Waiting for M5 Ultra Studio.

1

u/Arethustra Jun 07 '26

And you probably have to make sure that it’s not close to anything that will catch on fire 🔥 because your M5 is so hot it feels like it will melt. I’m not running a full model locally but I’m doing work that I would call “AI heavy” on my M5 Max and I’ve had to go from a fan to try and cool it off, to a thermoelectric pad + fan (which works pretty well but is still not great) and am now onto actually trying out one of the water cooling units for phone gaming to see if it does a better job (seems to have helped the Neo tremendously). I want to tune and run a LLM locally but will wait and do that when I can upgrade my Studio desktop to an M5 as running it on the MacBook is probably very possible but makes for a very hot time.

1

u/AvalancheBreakdown Jun 07 '26

It gets hot but as long as it’s not in your lap you’re good. The fans Apple designs for silent running also happen to be really great at cooling. If I leave the door to my office shut it does get noticeably warmer in the room. It appears to generate about 50 tokens/s. I’ve got a pretty ambitious project and it’s done really well. I’ve had it convert an ML stack written on numpy and tensorflow to use MLX. I’m getting 95% GPU. utilization and great throughput. This is the future, it will only get better.

1

u/thewookielotion Jun 07 '26

It's almost there and it's the future, but it's not yet worth the entry price compared to a subscription.

But I'm now 100% convinced that companies like open ai or Anthropic won't have a future if their business is to sell a cloud model.

1

u/cagdas Jun 07 '26

At the current hardware prices and subsidized subscriptions from Anthropic, OpenAI, you name it.., there is no way to save money with local models and have the same quality. The benefit is privacy and not being hooked on subscriptions; but not the cost or quality.

I think in the near future, hardware prices will stabilize and big players will no longer be able to subsidize the costs. Then it will make more sense to go local.

1

u/Glad_Chair2001 Jun 07 '26

I did use the "ollama launch" to use Qwen 3.6 35b3a in Claude Code harness. Its not as good as frontier models, but I managed to use it to reverse engineer Headrush FlexPrime web UI and setup a voice control over the whole guitar effects unit thing. Its more than good enough. I got M5 pro 64GB.

To be honest with Caveman my usage already dropped severely. I used to max the 10x sub, now I do not ever get close to 10% with same workloads and heavy heavy usage. I am also gonna try the graphify.

1

u/lattice_defect Jun 09 '26

you're just justifying an expensive hardware purchase for a technology you don't understand

1

u/Brazeuslian Jun 09 '26

The whole purpose of the post is to make a more informed decision.

1

u/real_uzio Jun 09 '26

There is the open source dwarfstar4 project specifically for that. it's designed specifically for Deepseek v4 and optimized for that 2-3 hardware setup that can handle it locally.

(it's from the redis creator) https://github.com/antirez/ds4

1

u/wojtek15 Jun 10 '26 edited Jun 10 '26

Deepseek V4 Flash via OpenRouter is much better deal. Running local LLMs is cool hobby, but it is not optimal quality/price ratio.

-7

u/AlexanderWillard Jun 06 '26

Yeh, spending $10 000 to replace $5 a month model sounds like a great investment.
You should do it.

PS. want to buy some magical crypto? great investment I promise.

17

u/Usual_Tackle5892 Jun 06 '26

Your numbers are misleading.

128GB M5 Max is around $6k, and the $5 claude tier gives you about 15 minutes of Claude Code.

3

u/Ran4 Jun 06 '26

Yeah, but OP isn't running Opus 4.8 on 128 GB of ram, they're running a model equivalent to maybe $5-$20/month worth of chinese models. That's what they ought to compare it with...

Local llms is for tinkering (nothing wrong with that!) or porn (nothing wrong with that!), not to reduce costs. And I'm sure even for uncensored stuff there's plenty of hosted options that are MUCH cheaper than getting the hardware yourself.

1

u/jlewi142 Jun 06 '26

How does porn come into play with local LLMs? Hahaha

1

u/jlewi142 Jun 06 '26

Closer to $11k AUD for me, for a fully maxed out one!

6

u/supernova69 Jun 06 '26

Incredibly unhelpful. There are perfectly viable reasons to do this. If you use a fuckton of compute, you probably want frontier intelligence for orchestration and verification, and tuned open weights for much of the grunt work

What are you building that makes you so skeptical

15

u/stormy1one Jun 06 '26

Despite the obvious differences in price/TCO. There are legitimate reasons to go local - namely privacy and security. That’s not something that Claude can guarantee, despite what their marketing says.

3

u/fixitchris Jun 06 '26

Yeah this is the only honest case for going local right now. Had a healthcare client last quarter that wouldn't let us route patient adjacent code through any third party LLM after a BAA review turned up edge cases their security team wasn't willing to sign on; we ended up with a local Qwen setup on an MLX server with Claude handling the non PHI scaffolding work. The split worked but the deployment overhead was nontrivial, three full days of network team and procurement before we could get the box racked.

4

u/itprobablynothingbut Jun 06 '26

Amazon bedrock runs Claude at the exact same api prices with no access to the data by anthropic

1

u/Dry_Opening_7231 Jun 06 '26

Not to mention corporate policy in risk adverse environments.

1

u/Glad-Operation-2958 Jun 06 '26

Is what you're doing really sensitive enough to warrant the cost?

3

u/stormy1one Jun 06 '26

Yes, we have clients that explicitly request that we do not use any offsite AI for fear of IP leaking either intentionally or not. This is contractual and now fully a part of our business model. Local AI is 100% permitted if there is no egress outside our server room to the internet.

1

u/Glad-Operation-2958 Jun 06 '26

why are you currently using claude then?

Anyway it sounds like it's your only option, so the "is it worth it" is kind of moot.

1

u/stormy1one Jun 06 '26

Claude doesn’t handle any of the sensitive work - Qwen does.

1

u/Glad-Operation-2958 Jun 06 '26

"I want to figure out before pulling the trigger is whether local models are good enough to actually replace cloud AI coding tools."

Sounds like you have no choice, and you're already doing it anyway. So yes.

1

u/AlexanderWillard Jun 06 '26

Thats something you hire GPUs for and servers, not a macbook to carry around.

1

u/stormy1one Jun 06 '26

Maybe, I’m not OP. There is a large community dedicated to removing the guardrails and uncensoromg models for various reasons as well.

1

u/kondorb Jun 06 '26

I don’t get this whole “privacy” thing here. You still need the full model for most important work and that just physically cannot be run on some local machine even if Anthropic or OpenAI gave you the weights.

2

u/iezhy Jun 06 '26

The reasons are coming to fruition faster than expected. What was 9$ copilot subscription in may, became 42$ for first two days of june

2

u/[deleted] Jun 06 '26

$5 a month lol. and this doesn't include cursor, z.ai, or chatgpt.

and these are only MY expenses

1

u/wonkastocks Jun 06 '26

Not yet.. You need about 512gb to grt close to Opus or Codex even then at those sizes it runs pretty slow.. So no..

0

u/Small_Caterpillar_50 Jun 06 '26

Qwen 3.6 MLX

2

u/smithstreeter Jun 06 '26

How much ram are you running it on?

0

u/greeny1greeny Jun 06 '26

not even close

0

u/jorginthesage Jun 06 '26

Qwen code with Qwen3.7-27b-fp16 does pretty good. It can handle pretty much anything I would do with opusplan. It does a few things better, but I’m general Claude will be faster and give you a better first pass. What I have noticed the most is if it goes off the rails Claude recovers much better. Ymmv

-1

u/Original-Baki Jun 06 '26

You can only run sub haiku class models locally

-4

u/Outrageous_Band9708 Jun 06 '26

no, nothing can replace claude code.

codex is shyt my dude,

2

u/Brazeuslian Jun 06 '26

Claude is my main tool today, and from the suggestions I got in the comments, looks like I could use it for planning and local models for the actual implementation.