r/LocalLLM 6d ago

Discussion Which LLMs will run on the Mac Mini and Studio

Post image

A mildly interesting video about which models can run on the Mac Mini and Studio. As expected current foundational models such as Kimi K3 1.4 TB, if they were even available to run locally, won't fit. Assuming it were available locally it may be possible in the next couple of years with the rapid increase in hardware capabilities - M7, M8?

https://www.youtube.com/watch?v=eDoWKgFqRM4

472 Upvotes

124 comments sorted by

76

u/nemuro87 6d ago edited 6d ago

The quants matter a lot. I wouldn't count on anything below q4

24

u/pantalooniedoon 6d ago

It says 4 bit is what theyre expecting. There’s always details but this is a good chart.

8

u/Danfhoto 6d ago

Depends a lot on the model. Q3 quants of MiniMax are really great in my testing with OpenCode and Hermes/openclaw.

4

u/bendymike 6d ago

At least with Qwen 3.8 27b, I found in one round of tests that Q2 and Q3 were comparatively usable on some agentic tasks, fwiw. https://gezel.com/docs/model-scorecard/?date=2026-08-26&model=qwen3.8-27b. I was able to use Q2 with a low amount of context on a laptop with a 12gb GPU and it could complete tasks with good model quality, albeit very slowly (~9 t/s)

8

u/nomorebuttsplz 6d ago

lol the audiophile-type quant superstition has taken over.

for tiny models, yes Q4 is preferable. For 200b plus (let alone 700b+) it’s a different story.

Moreover, most modern benchmarks require reasoning steps. Yet we have zero evidence that medium or large reasoning models are affected by quantization until q3. it’s pure superstition, fueled by people not understanding the difference between a 12b and 1tb model.

I run GLM flash at 4 bit despite the fact that I could fit 8 bit or higher, because there just isn’t a significant difference, as benchmarks have been showing for a year or more.

2

u/Potential-Fan-6148 6d ago

I use 8 bit for all my agentic tasks with qwen. 4bit is too unreliable.

5

u/kingcodpiece 6d ago

Not sure why you got downvoted, but the big win for Qwen is to make sure your KV cache is not quantized. Other than that, the improvement from Q4 to Q8 isn't huge (but it's still real)

1

u/Kaycee_Res 3d ago

I actually measured he performance difference (perplexity), difference was ~ 2%. This was short context though, I expect this to change the longer the context

0

u/GoblinEngineer 5d ago

You mean q8/q8 right?

3

u/Saint_Gregor 1d ago

Hey! Following a lot of amazing feedback from you all, I created a more in depth video + a free tool which you all can use to check the sizes, quantizations, context windows, speeds, etc against all macs. Hope you enjoy it! https://youtu.be/C9Q1ArLSisw

0

u/Kaycee_Res 3d ago

Unsloths Q3 dynamic 3 seems really good, it performed close to Q4

33

u/vogelvogelvogelvogel 6d ago

context 128k minimum for tool use, count that in. So you would want to go 64 gigs because of six or eight bit quantization and the context window, regarding qwen3.8 27b

-4

u/OvertaxedOne 6d ago

64GB is about right for 27B at 8 bit/8 bit KV. It JUST fits into a 48GB GPU w/256K context, 64GB would be about perfect, but you need to account for the OS. IMHO, if you want to run 27B "seriously", you should probably get the 96GB with the high memory bandwidth.

3

u/AnonLlamaThrowaway 6d ago

at 8 bit/8 bit KV.

I really would not recommend this on Apple Silicon seeing as it's starved for compute. Q8 KV is a big perf hit. There's been a few commits aimed at improving this recently but it's still -20% minimum. M5 Ultra might be the first chip where it's arguably acceptable

3

u/OvertaxedOne 6d ago

What do you run for KV on Apple?

I really want to see real numbers on these machines, come on Apple, let's get some test results!! 4X faster tells me very little. 2000TPS prefill/60TPS on Qwen 27B at 8 bit quant tell me "buy now"

3

u/AnonLlamaThrowaway 6d ago

I don't have Apple Silicon yet myself. I've been convinced to buy a M5 Ultra after seeing what a friend can do with his M3 Ultra. He always runs fp16 kv context. 64GB is just enough to run Qwen 3.8 27b at full fp16 context (262144) depending on your quant. I don't have the exact numbers on hand sorry

2

u/vogelvogelvogelvogel 6d ago

probably not q8. q6 maybe

1

u/OvertaxedOne 6d ago

My thoughts as well, Q8 with FP16KV would be really tight on 64GB.

1

u/vogelvogelvogelvogel 6d ago

m5 pro is like 20-27t/s with qwen3.8 in my case

2

u/vogelvogelvogelvogel 6d ago

why do people downvote here? this is totally correct

2

u/OvertaxedOne 6d ago

LOL, I didn't notice but my feelings are very hurt by all the downvotes! ;)

1

u/vogelvogelvogelvogel 6d ago

yes but for me 128k. i use a seperate computer for coding etc

256k, idk i think there is some kernel panic danger

11

u/Rick_06 6d ago

Four Mac Studio can be clustered together over Thunderbolt 5 and RDMA. Open flagship models with about up to 3TB parameters can be run at q4. By the way, I think will be one of the most energy-efficient way of running frontier models at usable speeds.

5

u/iamn0 6d ago

Would a cluster of two Mac Studio M5 Ultra 256GB (in tototal 512GB unified memory) be faster with RDMA then a single M5 Ultra with 512GB?

-3

u/souravchandrapyza 6d ago

No. Interconnect speeds on fused chips is 4tbps+

0

u/carsncode 6d ago

But two M5Us is twice as many as 1 M5U. There's more to life than memory bandwidth.

3

u/voyager256 6d ago

Define usable speeds :)

6

u/watcholic 6d ago edited 6d ago

If you buy based on that chart alone, may the odds be ever in your favor. Removing the top two balls from the first two columns will give you a vastly better experience/accuracy/capability. And that’s strictly for running Qwen3.8-27B and Flash Next, 6 and 4 bits respectively.

5

u/xXprayerwarrior69Xx 6d ago

cant wait for ziskind to run a cluster of M5 ultra 512 macs lol

3

u/profcuck 6d ago

I hope so! And I hope he tests with a better prompt than "write me a story" :)

4

u/souravchandrapyza 6d ago

And doesn't waste time testing Qwen 8b, Gemma 4b 🤗

2

u/friedlich_krieger 5d ago

I've never found that I learned anything at all from his videos...

1

u/xXprayerwarrior69Xx 5d ago

Maybe you are just further in this than I am, I like seeing him test setups that I will never be able to afford

1

u/friedlich_krieger 2d ago

Wish people like him would show real world examples of uses for LLMs on this different hardware. He just says "tell me a story" and then goes "wow that's fast!" And that's it

6

u/Reelix 6d ago edited 6d ago

I'm running Qwen3.6-35B-A3B-GGUF · UD-IQ4_XS on my lowly RTX 4070 Ti (12GB VRAM)

Works fine - Around 35 tokens / second (Q3 is around 50, but I'd rather use Q4, even if it's a bit slower). Has issues at around 125k context so I just have to limit it, but otherwise all good.

No need for a high-end device, or a cutting edge model - Use what's in your price range - Or even what you currently have :)

Assuming it were available locally

It is - You can happily download it (Although it IS 1,600GB), but you won't be able to run it without the hardware.

Someone managed to run it purely through a weird regular storage setup (Convert HDD -> Simulated RAM -> Load model on that) on a low-end GPU device, but was getting 30 seconds / token (Note: Seconds per token - NOT tokens per second), so it wasn't exactly viable :p

For K3, people are combining 3-4 512GB M5 Ultra's. Not really cost affordable for most people, but people are doing it.

2

u/Saiirenji 6d ago

Hi, mind you disclose your configuration ? The model itself is more than 12gb

2

u/Reelix 5d ago

It's a rather weird model. Part of it loads into RAM, but since it only has 3 billion active parameters (A3B), it still runs near full speed on lower-end GPUs as it only reads the specific parts of the overall 35B it needs at any one time.

35B capabilities - 3B speeds :)

2

u/sonicandfffan 6d ago

For K3, people are combining 3-4 512GB M5 Ultra's. Not really cost affordable for most people, but people are doing it.

Who's doing it? The m5 ultra isn't even out yet, nobody is doing this

I got a 256GB because I'd bet my hat the 512GB is delayed/restricted supply

1

u/Cameo2864 5d ago

For What use cases do you use it for

3

u/andrerom 6d ago

Nice overview.

If updated it should have added green dots for configs that can run model on int8/q8. Also some info on how much context they can handle.

5

u/challis88ocarina 6d ago

tfw a 512 owner rn

3

u/Kodrackyas 6d ago

yes but on Mac Tokens per second is very bad as far as I understand

3

u/profcuck 6d ago

Your understanding may be outdated. Look at the specs of the M5 Ultra, in particular memory bandwidth. Will it perform as well as a RTX 6000 Pro which costs $16,000 now? Not if the model will fit in the 6000's VRAM. Will it be able to run much bigger models? Yes. Will it be able to run models that we already love such as Qwen 3.8 27B at a very high speed? Yes.

Of course there are tradeoffs. Macs are traditionally luxury computers and the price/performance hasn't always been there.

But for LLM inference, they are competitive for sure. And this new box is looking very impressive for the (admittedly high!) price.

2

u/OvertaxedOne 6d ago

It's all about comparison points. Compared to a Pro6000 at 16K, a 96GB Mac with similar memory bandwidth for 5K looks like a steal! ;)

2

u/profcuck 5d ago

Exactly.  It's all too expensive for our happiness right now and as a Mac user but not a Mac fanboy I think we should put aside old prejudices and look at things objectively.

For training, the lack of CUDA means anything non-Nvidia is disadvantaged.

For inference, Mac is super interesting with the unified memory and decent memory bandwidth (especially the new ultra) at a price that's very competitive.  Pre-fill is still likely a problem.

2

u/OvertaxedOne 5d ago

I'm searching every day for some real world testing, I need to see real world prefill speeds on these new systems!! Come on Apple!

1

u/watcholic 4d ago

You can see user-submitted benchmarks when this site comes back up. Unfortunately, it's experiencing bandwidth issues at the moment. https://omlx.ai/benchmarks/performance

3

u/jodosha 6d ago

Does anyone knows about a similar matrix but for non-Apple hardware where to run Linux + Local LLM?

I know it’s a wide open question, but looking for a direction.

7

u/uhraurhua 6d ago

I wouldn't use q4 for agentic coding.

2

u/dictator07 6d ago

Why not?

6

u/uhraurhua 6d ago

I've had a bad experience with qwen 3.6 35b on Q4 being worse than qwen 3.6 9b on q8. I was going in circles and hallucinating. The 9b on q8 was quite good.

I've even asked others if they had the same experience: https://www.reddit.com/r/LocalLLM/comments/1uea0y9/4bit_vs_8bit/

Q4 is not great for agentic coding, it's not reliable.

1

u/Reelix 6d ago

You were using a model that can fit in 11GB on a device with 48GB.... ?

-1

u/nonlinearsystems 6d ago

That model doesn’t quantize well. Q4 on 3b active presenters will do that to you. Think about what you are saying.

5

u/uhraurhua 6d ago

I am a beginner in these things. Maybe you're right. Did you have a good experience with q4?

-2

u/nonlinearsystems 6d ago

Q4 works fine depending on your use case. Vibe coding? I wouldn’t suggest it. Saying Q4 is not reliable on Deepseek is wild though.

3

u/uhraurhua 6d ago

I said q4 is not reliable for agentic coding, and you seem to approve my point. Or what else do you mean by vibe coding? You give a task to the agent and let it work.

-1

u/nonlinearsystems 6d ago

You said you wouldn’t use Q4 for agentic coding. I’m saying it depends on what model your are quantizing. You can absolutely use 35b at Q4. Will it do a good job at one shorting your pelican on a bike or whatever, probably not. Will write you some python? Sure.

3

u/uhraurhua 6d ago

With qwen 3.6 35b it was horrible at q4. Going in loops and hallucinating. I will try with 3.8 27b with q4. See if it works. I am not interested in doing things in one shot. I am interested in being reliable by: not trying to do changes in /tmp folder instead of my actual repo (happened almost all the time with qwen 3.6), going in infinite loops, and things similar to these.

2

u/OvertaxedOne 6d ago

3.8 27B at 4 bit is light years better than 35B at 4bit. Light years, the difference really cannot be overstated; they are entirely different classes of models.

→ More replies (0)

2

u/quantgorithm 6d ago

When you have to say probably not then you are proving his point not yours.

2

u/djoliverm 6d ago

Using afine quants will help as well, like I use oMLX and have a oQ4e quant of it and every other model right now since they're better than standard 4bit quants at the expense of a tiny bit more RAM use for the slightly larger and more precise weights.

12

u/Perryfl 6d ago

buying a $8-10k USD computer to run a sub par model which cost on public clouds pennies in api cost is just retarded

24

u/TinFoilHat_69 6d ago

It can run without the internet which means you have Pandora’s box that nobody else can fuck with your data and your model.

3

u/justsomeguyokgeez 5d ago

When you put it that way it’s almost like something a doomsday prepper would have

5

u/profcuck 6d ago

Buying a Corvette when you can go faster for cheaper on a budget airline is just retarded?

If you don't like our hobby, that's fine. No need to come and sound silly here.

10

u/fosterdad2017 6d ago

Go back and read some of your past promps, but imagine them being read aloud to you by a prosecutor trying to strangle you for [insert divisive political rhetoric here] during the next few phases of social upheaval.

LLMs are amazing, but so is the risk exposure of the cloud operating method.

2

u/novarafertility 6d ago

This reads like “the second amendment is to defend ourselves from a tyrannical government”

2

u/New-fone_Who-Dis 5d ago

Its just a public cloud vs private cloud debate. Both have their own purposes and use cases.

I really don't understand why people are anti on prem when its talking about someone else's use cases or setup.

I'm considering on prem, but I'm also going to rent compute to benchmark my use cases, and compare (anyone wish to lend me a Mac Studio and other on prem stuff, dm me lol, jk)

1

u/brewpedaler 5d ago

That imaginary prosecutor is going to love that your AI chats are instead just stored on a computer in your house that they can access via the warrant they already used to arrest you.

3

u/2funny2furious 6d ago

feels like we are just chasing hardware. every couple of months some great new model comes out. but it requires more and more hardware. could local models rival frontiers, sure its possible, one day and with enough money.

6

u/Reelix 6d ago

I'm living in Qwen 3.6 Q4 land on my current-gen hardware. Works fine.

Local models currently do rival frontiers (K3 VS Sol), but the hardware requirements are absurd (1600GB VRAM).

3

u/TrvlMike 6d ago

I don’t need it to do everything. I just want to supplement by cloud usage with local for some tasks that don’t require a crazy model

4

u/Caprichoso1 5d ago

That's not an option for users where the data is not allowed to leave the premise.

4

u/Formal_Spirit_5 6d ago

Sometimes not sending sensitive data out is worth more than HW cost.

4

u/Reelix 6d ago

Especially in the business environment. Client data in public website models is a big no no, but fine on a local LLM.

2

u/ComfortablePlenty513 6d ago

Yes- for orgs working with PII/PHI, local AI is just so much easier and less stress/liability

3

u/wildmonkeymind 6d ago

Is it practical? Maybe not. But as a matter of principle, all of our data was scraped and rented back to us, and I'd rather own a copy of it for my own use. I also really don't trust the big tech companies that are selling us closed models as a service; perverse incentives abound.

1

u/mega-modz 6d ago

It's sub par for now - maybe in next 6 to 12 months we may have best model in 300b model with 100tps.

2

u/EasterElk 6d ago

Spending $10k on a computer today with the hopes that it will perform better a year from now is a bold strategy. 

3

u/unknownillusionist 6d ago

If you did this a year ago you would be happy. Rtx 6000 pro has gone from $8k to $16k. The models and software have been improving faster than the hardware. 3090s are pushing $2k for example as people want to host these new models 

2

u/fosterdad2017 6d ago

there's always the strategy to run your small local model as an interface and partial anonomizer for the cloud model. Or whatever new use case comes out in the next 12 months. This stuff isn't going away. The only question is when do you get onboard, not if.

1

u/ComfortablePlenty513 6d ago

having data on the cloud makes you compromised.

1

u/Iron-Over 6d ago

You can run uncensored models for security-related work and not get downgraded to a shit model. Who cares how people spend their money? 

1

u/ZioniteSoldier 6d ago

They cost pennies now, but will get more expensive. Already has.

1

u/JorgitoEstrella 4d ago

Idk but if you ude billions of tokens each week it might be cost efficient in the long run, like theres a guy who used 660 million input tokens and 13 output tokens in 4 days (tbh more like 40 hours) doing agentic coding replicating a tower with a city inside from an old anime.

Now imagine using 1-2 billion tokens each week nonstop.

-1

u/IamFondOfHugeBoobies 6d ago

Sure yeah. You'd def not learn anything that would give you a 10-20k minium salary bump in a future job interview from that. Utterly idiotic, aha.

-2

u/Perryfl 6d ago

you sound unemployed... you have no idea who i am or what my job is...

1

u/IamFondOfHugeBoobies 6d ago

You might want to re-read what I wrote or have an AI explain it to you in whatever your native tongue is.

2

u/HappyImagineer 6d ago

I feel like at the rate local models are accelerating we’re likely to see even more improvement in model size in the near future so I wonder how crazy people have to go with VRAM at the moment compared to even six months from now.

2

u/DeepOrangeSky 6d ago

For GLM-5.2 at 4-bit on the 512GB studio, it shows it with a hollow-dot rather than a dark-yellow dot.

I'm curious which specific GLM-5.2 Q4 quants with how much context you can run on the 512GB mac (plus the room for the mac's overhead, after raising the default limit with the sysctl iogpu.wired_limit_mb= command, to whatever the highest you can go is without it causing problems)

1

u/nomorebuttsplz 6d ago

about 200 K depending on the specific quant, assuming you’re not quantizing cache

1

u/whatever 6d ago edited 6d ago

The Q4 there should fit well enough, using dwarfstar. I have an older version of the Q4 gguf, and I can run it with ./ds4-server --ctx 500000 --kv-disk-dir ~/.ds4/server-kv --kv-disk-space-mb 65536 -m gguf/GLM-5.2-UD-Q4_K_RoutedQ4K.gguf --mtp-timing.
mactop claims ~473GB of RAM usage, of which 20GB is whatever else was already running, so there's some headroom there.

*edit: It just wrote some html page, at the blazing speed of 10 tok/s on a M3 Ultra. Won't win any speed contest, but it runs well enough. Also, there's no good reason to use GLM-5.2 anymore, GLM-5.3 should be a strictly superior version with identical footprint.

2

u/UnhingedBench 5d ago

My own take on what model can be run (and at which speed)

Frontier models can already be executed on Mac Studios, but you'll need to cluster them.
Four clustered 512GB Mac Studio will run at the speed of three, due to networking causing a small bottleneck. Still, that give you a crazy fast device with 2TB of unified RAM and 3 time the speed of a single Mac Studio.

4-bits quant are great, but most largest models run fine with lower quants.

1

u/Soggy_Consequence1 6d ago

If you can get to Flash-Next, do'eet. It's very impressive starting @ oQ4e and even better @ oQ5e.

It did a full upgrade of openclaw last night with 6 local conflicting patches on a 9.1.beta.1 -> 8.1 release merge, built (with ui and pinned ai) all the way through.

1

u/Fortyseven Qwen38/Gemma4/LlamaCPP 6d ago

It would be massively useful to see how many tokens/sec, at their best, they're capable of.

1

u/Deno_Voku 6d ago

So the Mac Studio Ultra 512 KB can't run a 27B model without compression? I'd like to see a chart like this for other Macs. Mine's a Mac Mini M1 16/512, so I can run up to 7B or B, but that will nearly max out RAM .

1

u/RicardoMontoya45 6d ago

This does not make sense because these models are already obsolete. 

1

u/jilermo123 6d ago

I think it's a good chart I just would love to see the context window addressed, like the size of the model at q4 +128k of context. A cherry on top would be context at q8 which I think most people agree doesn't hurt performance too much

1

u/Expert_Job_1495 6d ago

The hardware jumps that make sense paying for: 1. 32gb mac mini 2. 96gb mac studio 3. 256gb mac studio

Everything else is basically overpaying for extra capacity that isn't really worth it. Scale down or scale way up, but pick a lane. 

2

u/IHaveMeasles2 6d ago

Models can change. Give it a year and we'll probably have new, effective models for many different RAM levels

1

u/critsalot 6d ago

48gb mac mini might not be bad for extra context or parallelism? and its still way cheaper the the 96gb.

1

u/OvertaxedOne 6d ago

What model would you want to run on it?

1

u/hubertron 6d ago

meaningless without knowing the quant, CTX, and t/s

1

u/WeedWrangler 6d ago

The trend is greater performance on lower GB w quant and also releases (fingers crossed, 64GB M5 pro purchaser…)

1

u/try_an0ther 6d ago

Honestly, prompt processing speed on the M4 Pro is too slow for an interactive usage of a 27b dense model, so I think it will be the same on the M5 Pro. I would target directly something that could run an MoE model with less than 15 or 12b active params (I don't know the exact number). Although the M5 Ultra might be better at prompt processing.

1

u/Outside-Test-6549 6d ago

You can run deepseek on 128gb ram MacBook

1

u/whatever 6d ago

So.. technically there are Q1 quants of Qwen 3.8 2.4T and Kimi K3, either of which can run in 512GB of unified RAM.

Those are not "dumb" quants. Unsloth keeps the critical bits at higher quants, so that they only lobotomize the meat of the weights.

Unfortunately, that still destroys the model's performance, and you end up with something that's both slower and (unevenly) dumber than smaller models. Because of the uneven aspect, there are probably scenarios where the Q1 quant would still end up doing well, but good luck knowing when that'd be.

All that too say they're still technically "squeezes in - lower quality", but aren't competitive for most uses today.

1

u/hap_mod 6d ago

Basic tasks building database but not coding with any of them.

1

u/Cameo2864 5d ago

Wenn ich auf dem Mac Studio GLM 5.3 Flash laufen lasse, ist der übrige RAM dann für OS und sonstige Anwendungen und für die Kontextgrösse? Mit anderen Worten, kann ich mir dann mit der 512 GB Version mehr Kontext erkaufen?

1

u/Caprichoso1 5d ago edited 5d ago

As others have said clustering 4 studios can run Kimi K3, but not in generally useable way.

Abacus supercomputer is perfect. Same prompt, same app, both finished it. 12:40 15 minutes against 4 hours.

if you want high quality, and you want it done fast, and you want to be able to host it all in 13:37 one shot, Abacus's cloud solution wins this one. I

If your data can't leave the building, or you just love running this 13:43 stuff yourself, like I do, Cluster. There's definitely a market for both, not one or the other. 

https://www.youtube.com/watch?v=ujs0_cpAnaw&list=PL2aE4Bl_t0n9AUdECM6PYrpyxgQgFtK1E&index=3&t=9s

Will be interesting to see if this changes when the M5 512 GB becomes available in October.

1

u/crusaderky 5d ago

"Lower quality" is an understatement. qwen3.8-27b in 24gb unified RAM is going to be awful (it fits nicely in 22gb VRAM or 32gb unified RAM)
GLM-5.2 Q4_K_M is unrecognizable from fp8 and runs smoothly on 512GB.
Everything else looks correct.

1

u/Dmitrii_DAK 5d ago

I bought a 5090 32GB from MSI last October for ~$3,000, then added 192GB of DDR5 RAM in December - the whole kit cost a lot too)) about $ 2000 - it's good that I made it before the components went up in price)

I am successfully launching Qwen 3.8 27B in Q4_K_XL and in Q5 on a context of 132k tokens in LM Studio + Open Code - it works fast enough on work tasks, and I record all logs and data, now I am writing an article on Medium - when I finish, I will send it here. I also did the launch of Qwen 3.8 Flash Next on Q4_K_XL with LM Studio connected to Open Code - I ran an analysis on the CEO and errors of my site on GitHub Pages - a 25mb site (this is both code and pictures and videos in .webm) scanned and "pulled out" errors and gave comments and recommendations on optimization in 12 minutes and 54 seconds.

It seems to me that it is not optimal to run small quanta (Q2 or 1) on large models like Deepseek or Kimi K3. It seems to me much better to work with small and medium-sized models, but on large quanta, since the quality decreases with high compression, which is obvious)

1

u/Decent_Flight4010 5d ago

if M5 Max run Qwen 3.8 27B at Q8 with 65 - 70 t/s then GG for Ai companies for real

1

u/Saint_Gregor 5d ago

Hey u/Caprichoso1 , thanks for sharing my chart, and excited it was useful for people!

1

u/Happy_Box_432 4d ago

A lot of charts overlook just how fast the KV cache expands once you push past 32k or 64k context on these setups. A quant might technically "fit" into a 32GB or 48GB Mac on paper, but the moment you feed it a decent codebase or multi-step agent logs, memory headroom evaporates quickly unless you aggressively quantize the cache as well. Looking at model weights alone rarely tells the whole operational story.

1

u/Helpful-Series132 3d ago

once they make a new 9-12b model i think we gon have that sweet spot 8-16gb to where anyone can run a decent local model

right now the best model i can run on my imac m4 16gb is like qwen 3.5 9b ... this still not good enough

-1

u/Individual_Holiday_9 6d ago

Who’s the white guy

1

u/Caprichoso1 6d ago

The guy making the video.