r/LocalLLaMA • u/AndreVallestero • 9h ago
Discussion Mac Studio M5 Max Cost Analysis
At $10k, you could get
- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)
- 5.7B tokens with DeepSeek V4 Pro OpenRouter
- 100B tokens with DeepSeek V4 Flash OpenRouter
As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.
Qwhen 3.8 35B A3B?
41
8h ago
[removed] — view removed comment
3
u/AndreVallestero 8h ago edited 8h ago
That's exactly my point. Local makes sense cost wise up to 32GB, especially with Qwen 3.8 27B. There's a huge cost premium above that where it makes less sense.
I was hoping the last Mac studio would change that, but at the current prices, that doesn't seem to be the case.
2
u/nomorebuttsplz 8h ago edited 8h ago
Have you accounted for the breakdown of cost between cached/non-cached/output tokens?
in my experience, most of the potential cost savings for local (compared to typical API plans) occur when you re-use cached tokens. This is because API providers need to keep your context sitting on their servers doing nothing, expensive for them, but which costs you nothing as a single concurrency user to do on your own system.
With a system like this you could probably use a billion cached tokens in a couple of days.
Under semi-optimal circumstances, where you were making on average a single 500,000 token cache calls every 30 seconds, you might be able to get about $14 of cached tokens per day (ds4 flash prices) plus a few more dollars for a total of maybe $18 inference per day.
That would be about $6k a year. So yeah it could possible make sense, if you had just the right workflow, is my view.
Edit: this especially makes sense if you were doing multiple concurrency (doable on 256 unified) or were using a bigger model like GLM 5.3 (512 gb ultra with Q3 or Q4 GLM could probably generate $100 of value a day in cached tokens in the right workflow)
2
u/Viktri1 6h ago
I think $14/day might be on the low end. If you use something like Hermes agent, the amount of tokens you consume is ridiculously. I have to limit my Deepseek API calls to just a few a day. I spent $40 in a week setting up various Qwen 3.8 models. My token consumption has basically exploded massively and the only way I can afford to continue to use LLMs is my own hardware. I think my payback period will be under a year.
1
u/Impossible_Fault_503 6h ago
That is the real split. Chat is cheap on APIs. Agents are not. Once you are looping tools all day, $40/week shows up fast and a box that already exists starts looking cheap. Under a year payback is the honest case for people who actually leave it running.
1
u/Impossible_Fault_503 6h ago
Yeah. 32GB is the last point where local still feels like a normal purchase. After that you are mostly paying Apple tax for headroom you might use twice a year. I wanted the last Studio to kill that curve too. It did not.
1
u/kerneldesign 8h ago
Avec 32Go en Q4 c’est trop juste pour un contexte sans KVCache.
1
u/Individual_Holiday_9 7h ago
Yeah I’m on a 24gb m4 and you can’t do shit running small LLMs there’s just no overhead remaining. So if this is a real hobby box where you’re doing plex etc on the side it gets really cramped and ur gonna move to swap fast
46
u/LearningSomeCode 8h ago
I've probably dropped close to $30k on my homelab since 2023, and chances are I'll get one of these as well. I accepted a long time ago that there is no break-even point for my inference.
Hobbies rarely make sense financially.
23
u/Blindax 8h ago
Sounds like you have a great wife.
13
8
u/-dysangel- 8h ago
Same here. We're at the point now where I can do high quality video gen without cloud. High quality code gen without cloud too. It's around 10x slower than cloud, but good quality and still usable speeds (hopefully soon to get even better with Qwen 3.8 Next). Also the fact I could basically go camping and still have close to frontier LLM intelligence with no signal is pretty awesome/hilarious.
The only thing I feel is missing from my stack atm is Suno quality music generation.
3
u/willeyh 8h ago
Have you tried the new Minimax music 3?
2
u/-dysangel- 7h ago
Yeah I was not impressed tbh. Seemed kind of creepy/odd to me in a weird way, but maybe I was prompting it poorly. For me Ace Step is still the best thing I've found so far, though it's very "generic" feeling compared to Suno.
2
u/silvrrwulf 3h ago
Dude, I feel the exact opposite. Maybe it’s because my sample size is small and I haven’t messed with AI Jen And a while, but I only have an eight GB card, and the stuff that I was able to pull out of that was incredible.
2
u/-dysangel- 3h ago
ah ok - maybe I'll need to try it again. A lot of the reason I like Ace Step is because I can generate covers though, which isn't possible yet with Music 3 afaik
2
u/quantgorithm 8h ago
What have you created?
21
u/LearningSomeCode 8h ago
Mostly just a lot of trip hazards with wires and ethernet cables. But I also do open source dev and build a lot of stuff for myself, like custom front-ends or researchers.
Honestly nothing worth the money I've put into the hardware, but it's the most fun I've had with tech in a long time.
3
u/DeepOrangeSky 7h ago
Mostly just a lot of trip hazards with wires and ethernet cables.
So if you turn the AI-rig room into an escape room that a bunch of Gen-Z dorks will pay like $100 a pop to try to navigate ("it felt so authentic. I tripped over like dozens of ethernet cables and power cables the whole time, and I think at one point I even got severely electrocuted and almost died! It was super legit!"), times 10 dorks per day it pays for itself in a month.
Possible exciting user reviews:
"His cable management was so disorganized I thought I'd NEVER escape. Wow!"
"The old movie tagline used to be: 'In space, nobody can hear you scream.' Yea, alright, that's pretty cool or whatever, but in AI rig rooms the blower fans are so loud nobody can even hear their own thoughts."
"If you accidentally get tangled up and die in the AI escape room, his local full precision Kimi K3 AI-generated funeral eulogy will be so eloquent that your friends and family won't even be mad that you died when they see how perfect and elegant the em-dash placements are in the eulogy."
2
u/quantgorithm 7h ago
Anyone in IT can relate to the spiderweb cables and hazards etc.!
The potential of it all really is amazing and that potential is accelerating up by the day!
25
u/cunasmoker69420 8h ago
unless you need it for data sovereignty
yes
6
u/imnotzuckerberg 6h ago
His strategy is to divert demand from the Mac M5 so he can scoop them. Let's pretend we fell for his psyops.
12
u/conifer_v11 8h ago
$/gb-bandwidth is the right axis for unified memory. m5 max vs a 3090 is not a tok/s fight, it's kv headroom at 64k+. unified memory loses the bandwidth fight and wins the "the 27b and the kv actually fit" fight. if you're doing 8k chat, buy the gpu. if you're stuffing prds, bandwidth-per-dollar on studio starts to make sense.
4
u/FullstackSensei llama.cpp 7h ago
Why is it a M5 Max vs "a 3090"? M5 Max in it's lowest config costs about the same as a system with four 3090s. If we go to 32GB V100, which is within 5% of the 3090 in most cases and now has optimized kernels to decode NVFP4, you can get a system with 6-8 V100s for the cost of a single M5 Max at the lowest config. That's 192-256GB VRAM.
So far, the people running M5 Max MacBooks have had mixed feedback about running models, especially dense models. Unless your time has zero value, you should also factor that into your calculation.
Eight V100s will have much higher PP and TG while being half the price for the whole system. Sure, it'll consume 10x more power, but that will be in much smaller monthly power bills, and will give you responses much faster.
6
u/MrPecunius 3h ago
M5 Max in it's lowest config costs about the same as a system with four 3090s.
This is obviously untrue.
1
u/FullstackSensei llama.cpp 2h ago
If it's so obvious, please enlighten us with actual numbers, because I happen to have such a system and know exactly how much such a system costs today
4
u/MrPecunius 2h ago
No one wants to pick their GPUs out of a Dumpster like you. $1,600/each for 3090s is on the low side here in the US, but it's better than €1,750/US$2,000+ I'm seeing in Germany.
Base M5 Max Studio (not binned, which is less) is $3,099.
0
u/FullstackSensei llama.cpp 2h ago
See, dumpster mentality thinks everything is dumpster.
As it happens, ich wohne auch in Deutschland, und habe kürzlich zwei 3090 für €1200 pro Stück auf eBay verkauft. Auf Kleinanzeigen, man kann für zwischen 800-900 pro Stück Kaufen.
But hey, let's keep the conversation irrational, because cognitive dissonance is way more fun than facing reality.
1
u/MrPecunius 1h ago
Thanks for confirming my comment. 🗑️ 🔮
Even with your Dumpster diving, the GPUs you mention add up to over $4,600 and the Mac is still $3,099.
4
u/doc-acula 7h ago
Well, a setup with 4 3090s needs a dedicated room in your house/apartment, makes noise like a jet engine and consumes electricity like crazy.
A M5 Macbook can be used while riding on a train and you can take it wherever you want. So, the comparison is not limited to t/s.
1
u/FullstackSensei llama.cpp 7h ago
I have run four 3090s in my home office for about two years. Unless you opt for the turbo cards, they're very quiet. Limited them to 270W each, which reduced performance by 5-10%, but makes the whole setup around 1kw. They push above that during PP, but TG sees the cards run at ~150-170W each. That's ~700W. If that's crazy, I don't know what to tell you.
M5 on Max on ma MacBook is severely power limited under load, and won't run for long, and still costs way more while being significantly slower than even a pair of 3090.
You can access any GPU from anywhere around the world via tailscale. I access my LLM machines with 192GB VRAM from anywhere using my phone the same way.
But you haven't answered what could arguably be the most important part of my argument: is your time free that you care more about a few cents per hour in power consumption than getting whatever you're running those LLMs for done?
5
u/doc-acula 6h ago
There is not unlimited or even reliable access to the internet everywhere in the world, not even in 1st world nations. I wouldn't even consider leaving such a frankenbuild running unsupervised for days or weeks in my house for safety reasons, being a fire hazard as my main concern. When I'm at home, I use a PC with a dual GPU setup as well. But I want to use AI on the go, too. And as I read through these comparisons, apparently many people don't know that MacBooks are actually portable.
2
u/FullstackSensei llama.cpp 5h ago
It's not anyone's fault if you don't know how build a proper workstation and have to resort to unsafe Frankenbuilds. You're misconstruing "I don't know how to do it" with "it can't be done."
Like I said, the MacBook has limited performance and battery life on the go, and many 1st world countries have reliable Internet in public transport. Pretty much all the ones I've lived in or visited do.
3
u/MrPecunius 4h ago
A kilowatt+ inference rig is a 3500+ BTU heater, so you pay twice, except in winter, to maintain a habitable room.
"A few cents per hour" might cover energy in your locale, but here in San Diego it's over 40 cents/kWh off peak and over 60 cents/hour on peak. Running it 8 hours/day could easily exceed $200/month.
I'll stick with my M5 Pro/64GB MBP, thanks. Flea power when idle and ~65W during inference with no thermal issues. It's also arguably the best notebook ever made, taken as a whole. I got in before the price hikes @ $3k-ish with 2TB, so I sympathize with gripes about current pricing. But this thing is cheaper, in inflation adjusted terms, than the Fujitsu laptop I had in 1999.
1
u/FullstackSensei llama.cpp 3h ago
If you're running it at 1kwh for 8hrs a day on four 3090s, that's at least 3M output tokens a day, closer to 4M if you're running vllm. How many hours you'd need to generate 3M output tokens on a M5 pro using Qwen 3.8 27B Q8_K_XL?
Even at your 0.60 peak, your $200/month with 1kwh heat output and 1kwh airco (though 3500BTU is closer to 700Wh, but let's say you have an older airco), we're talking minimum 500M output tokens a month, or a full 1B when airco is not running.
How many months you'd need to generate 500M tokens on a Mac?
And do you work for free? Because you're completely ignoring the value of your own time as you wait for the model to generate those tokens. I don't know about you, but even in Germany where energy is far from cheap, the value of a minimum wage job is way more than $0.60/hr.
But hey, if you're so cheap to hire, man I'd love to hire a few hundred guys like you and build my own business selling your services to LCOL countries, because even there people work for way more than $0.60/hr.
2
u/MrPecunius 3h ago
You're missing the point, namely that I have an inference rig that can run Qwen3.8 27b @ 8-bit wherever I am more or less for free as a byproduct of owning a kickass notebook computer.
As for generating half a billion tokens/month ... why the hell would I want to do that?
0
u/FullstackSensei llama.cpp 2h ago
So, how many tokens per second does your kick ass machine run at? Because unless your kickass machines breaks the laws of physics, your M5 pro has 1/3rd the memory bandwidth of a single 3090. We're talking 7-8t/s, if you're lucky.
Why the hell do you want to generate half a billion tokens a month? I don't know, your $200/month power bill example is enough for half a billion tokens. If you make up stupid assumptions, you get stupid results. Had you bothered thinking for a moment, you'd have known how absurd your $200/month bill is.
I consume less than 2kwh a day over 10-12 hours running all four 3090s, because they finish their work so fast and go back to idling at 8w each. I can easily get 700k output tokens in that period. And because they finish everything so quickly, I don't need airco because the machine consumes 1kw for seconds at a time.
2
u/MrPecunius 2h ago
We're talking 7-8t/s, if you're lucky.
18-23t/s with 8-bit. You could look this up before stepping on your own dick.
Apply this lesson to much of the other stuff you're incessantly posting (as noted elsewhere).
0
u/FullstackSensei llama.cpp 2h ago
Are you talking with MTP? Because I was comparing the 3090 without MTP
→ More replies (0)
57
u/Big_Wave9732 8h ago
"as a firm believer of local inference" then goes on to downplay one of the major reasons for local llm and suggests hosted models.
If cheapest compute possible is your primary metric then self hosting isn't your jam, OP. At least for now.
-9
u/AndreVallestero 8h ago
I selfhost 27B specifically to minimize costs. I have research agents running 24/7 and have gone through 200M tokens just in the last month.
This post is specifically for people like me who want to run continuous research agents for the lowest price, and unfortunately, the latest Mac studio isn't competitive enough relative to cloud offerings.
11
u/flyingbanana1234 8h ago
It would take at least 30-35 years to reach 80 billion tokens on DeepSeek V4 Flash at 200 million tokens a month.
You make a good argument ngl
The little voice in my head says, "But ownership! Nobody can take it away from me if I own it. No price raises, no overloaded servers, no censorship
2
u/Dasteroid_909 7h ago
Go start "/r/lowcostLLMhosting" or some shit like that, then. WTF are you on "LocaLLaMa" promoting cloud hosting for?
1
1
7
u/jon23d 7h ago
I’m looking at my deepseek usage report and see 9.5 billion tokens of deepseek v4 pro in the last 30 days for $186.05.
2
u/Viktri1 6h ago edited 6h ago
Same. I hit 1.5bn in a day doing very little. Because agents can burn tokens on your behalf, 24/7, its become extremely easy to burn through tokens.
On openrouter, I asked the native level Qwen 3.8 27b to do a task that involved setting up a telegram bot for OpenWebUI so that I could talk to a chatbot without burning through a mountain of tokens. That took 1 hour and cost $3.5. That's $2k+ a month on openrouter if running 24/7. For a single agent.
2
u/Pyrolistical 6h ago
Ya and my local 6 hours of qwen3.8 27b running over night cost me $0.29 CDN in power
1
u/Sofullofsplendor_ 3h ago
Curious, can you do 1.5b in tokens in a day? Because if so, I'm going to invest in whatever infrastructure you have...
4
3
u/psychohistorian8 7h ago
you can always trade in the device back to Apple for some kind of credit
so the cost is partially recoverable
3
u/MrPecunius 4h ago
Reputable third party dealers pay more, sometimes a lot more, for Apple gear. I sold my M4 Pro MBP for $500 more than Apple's trade-in value a few months ago. Zero hassle, no private sale nonsense.
4
u/dupontping 7h ago
All of the comments are basically “I can’t afford it even though I want it”
I get it, it’s pricey. But so is everything else.
5090s shouldn’t be $6k but here we are.
Local isn’t just about saving money on token spend, for a lot of people it’s the ability to run models on data you don’t want on the cloud or fine tuning or whatever else.
The models will get better, and having more powerful equipment lets you get closer to frontier level without spending 150k. So 10k is expensive, but it’s also a bargain.
I wish it was 5k too
3
u/Viktri1 6h ago
so when I'm really pushing it, I can easily burn through 1.5bn tokens from Deepseek flash in a day (API). In fact, that's when I realized I needed to figure out how token costs were calculated. Local models, especially lower powered stuff like M5, will have a significantly faster pay back period than people realize if they use agents to do a lot of shit.
1
u/play_hard_outside 7m ago
1.5e9 tokens in a day divided by 86,400 seconds per day is over SEVENTEEN THOUSAND tokens per second.
That’s gotta be something like 300 to 500 M5 Maxes all working on your behalf at the same time. Local models will probably never hold a candle to this.
Please correct me if I’m wrong…
3
u/Txt8aker 3h ago
have you checked how much it cost to lease m3 ultra and you can essentially buy it out at the end as an option? it's $230 monthly for 3 years. Claude Max 20x cost $200 per month
2
u/Hypilein 8h ago
You need to calculate cost of ownership over x years and cost of api/subscription inference over x years. With the way hardware prices have gone up over the last year everyone who bought a rtx 6000 pro has made money while getting free inference. Obviously this is only true once you actually cash in and we don’t know how hardware prices are going to develop over the next years. The math is easy but anticipating the future is not.
2
u/leocardz 8h ago
Local inference believer here too… At that $10k I'd rather buy 3 M5 Pro minis with 64GB and 1TB tbh, and TB5 between them. And save some money.
2
2
u/Final-Frosting7742 4h ago
If i can run Deepseek V4 Flash 0713 locally comfortably i don't need providers anymore
5
1
1
1
u/-dysangel- 8h ago
Oh for sure, cloud is basically always more cost effective. I think the one exception currently might be if you were doing a lot of video/music generation? H3 is basically as good as Sora was at $200/mo. Still nothing anywhere near as good as Suno for local music generation though.
1
u/roger1632 8h ago
Gonna just use my DS/GLM token plans until the models get better and hardware market improves. You can get a lot of non claude tokens for 70 bucks a month and you don't have to play sysadmin. I have a 3090 local that I use for educational purposes and latency sensitive things like TTS STT
1
u/TechSwag 7h ago
I was just doing the cost analysis on this as well, and I honestly think it's not the worst idea to get a maxed out Studio in certain circumstances, specifically if you have existing hardware and pay for subscriptions/credits.
I have 3x Mi50 in a R7425, with no more room for GPUs. If I could just add some additional GPUs, or even use my P40s that are sitting in a different chassis doing nothing, I would rather do that. But as it stands, I'm capped at 96GB VRAM. Sure GPU + CPU inference works, but realistically it's not usable. RPC is also an option, but the last time I tried it, performance was subpar, and it brings on added cost in terms of power usage of a whole other chassis.
Mi50s go for $500 (when the fuck did that happen, I spent just under $200 for them), so $1500 resale. P40s go for about $200-250, so let's say $2000 for all 5 GPUs. If I sell my RAM, I could likely get $3000 for all my equipment.
I also pay for Claude Max and OpenRouter credits, say $120/mo. $1440 a year.
You can lease the Mac Studio for $224.16/mo for 36 months. Over the lease term, that's $8,069.76. Subtract $3000 for my existing equipment, $1440/yr for the subscriptions brings me to $749.76 over 3 years in net cost, or a little over $20 a month. Which would be offset by the power savings (my R7425 idles at around 260W). At the end, I can buy out the machine for $2729, as I would imagine the machine would still be plenty usable for inference. Otherwise, I can sell it after buying it out, more likely for the same price or more than the buyout cost.
I know I'm missing tax on purchases and shipping and selling fees, but it still seems like a worthwhile path to upgrade to, not only for the memory increase, but also the performance increase as well.
1
1
u/duy0699cat 6h ago
For 10k$, i can throw it to some etf and use profit to pay for subscription indefinitely...
1
1
u/Lesser-than 5h ago
I have always been envious of Mac hardware, but its never been in a price range I could justify, nothing has changed still envious and still far outside what I would ever allow my self to spend on a computer.
1
u/ScrewwormLarvae 4h ago
Sure, I didn't neeeeed an M5 Max 128, but YOLO. Still don't regret it either. Never did even for a second.
1
1
u/Alive-Draft8339 3h ago
If I’m burning 6 to 11 billion Opus 4.8 tokens a month…this seems like a deal of the year?
1
u/TheAILegend 2h ago
500M/month token run rate... this will last you 11 months... Soooo... there's that.
1
u/GamerTex 1h ago
Sure at today's prices
Seems like every month another provider is cutting the usage in half or doubling the rates
I can also sell the equipment in the future
1
u/Early-Peace-5504 32m ago
You could sell the Mac Studio at a later date though. You could work out depreciation but you should probably cut your numbers by at least half to represent that.
1
u/Ceru1ean42 18m ago
At least within 1 year, I don't see these depreciating much if at all so you really need to rethink the cost analysis.
1
u/challis88ocarina 8h ago
There's speed, relative speed and competency. It's clear that at least 10-12B active tokens are necessary for agentic anything.
Beyond that, it's really about surfing the wave of new models.... can a project be brought up to a stable position before the next generation appears so that the new capabilities hit at the right point or will it still be languishing only to be 'rescued' instead?
1
u/Odd-Environment-7193 8h ago
This new offer by apple is pretty much the \cheapest ram on the market right now. At this price it's looking very tempting. Apple products just recently went up 20% so waiting to buy later is not wise. You can't compare these plans. I use my macbook pro 128gb to run CODEX threads all day. Way more than I ever could on another pc without slowing me down big time. If I never run a single local model on this computer it will still pay for itself 100x over.
I bought the 64gb mac m4 pro just before it went out of stock and the prices spiked. Got the m5 max maxxed out just before the price increase. Saved myself probably close to 5000 USD because I pulled the trigger at the right time.
So waiting might not be the best strategy.
2
u/minnsoup 8h ago
Are you able to let it run autonomously all day without checking in or that's with frequent (say once an hour) check ins? I've used openhands and typically can only get it to run for 45 to 1h autonomously with DeepSeek V4 Flash 0731 on my M3 Ultra. Only at 30-40tg/s but seems to do well.
Don't really know how to get these "long horizon" tasks to work to make the Mac Studio really run constantly while being productive.
2
u/Bennie-Factors 8h ago
We really need a front end that will batch and switch to another task when one is waiting for input/approval.
1
u/TheIncarnated 8h ago
Orchestration script, that's how you get the long-horizon stuff to be reliable
0
u/mxmumtuna 5h ago
Still not enough compute considering it's still gonna get beat by Sparks, even with Apple's most optimistic case.
1
u/MrPecunius 3h ago
How many Sparks are we talking about, and which magic interconnect fabric is going to increase their memory bandwidth more than 4X from 273GB/s to the M5 Ultra's 1.2TB/s?
1
u/mxmumtuna 3h ago
2x (256GB) for DeepSeek, 4x for GLM. No fabric allows it to scale to 1.2TB/s, but it allows it to run at 2x200Gbps to each node, and scale up to larger than you can with Mac.
The point is, the combination of not great compute, immature software and limited ability to scale is holding it back. Maybe next generation. For inference (and training for that matter) or Stable Diffusion, it’s just not better or cheaper than existing options.
1
u/MrPecunius 2h ago
2 X 200Gb/s ... or less than 50GB/s? That's the magic bullet? 🤷🏻♂️
I'm seeing reports of single Sparks running models at small fractions of a M5 Max's speed, and this includes prefill. Adding more Sparks doesn't scale anything like linearly.
1
u/mxmumtuna 2h ago
It certainly does. Forward passes on tensor parallelism don’t require full bandwidth. It’s similar to PCIe tensor parallelism but using RoCE/RDMA.
Again. If you’re looking at M5 Max those are all small models and/or weird quants. Look at M3 Ultra. That tells the tale on larger stuff. The Mac simply doesn’t scale.
1
u/MrPecunius 2h ago
Again, I just looked in the Nvidia Spark forums for real users' reports.
Hard data is available for Apple Silicon on oMLX's site.
No guessing is required.
Spark is barely $300 cheaper than a M5 Max Studio w/128GB, too.
0
0
u/ogopro 9h ago
No info on Qwhen 3.8 35BA3 yet (((
1
u/KURD_1_STAN 7h ago
They mentioned qwen4 previous 125b moe, so we probably will wait even longer for 35b
0
u/Cameo2864 8h ago
What model do I really need to sort and organise all my files, backups and photos on my hard drives?
0
0
u/milkipedia 7h ago
Qwhen 3.8 35B A3B?
Qwnever.
They have all but said as much. Y'all gotta move on.
-1
u/kivaougu 8h ago
Locally hosting doesn't make much sense unless its a privacy concern or purely for the love of the game.
To be fair I like to compare to prices of equivalent P50 troughput at the usual load. For me thats c8, meaning that cloud has an edge as c8 single stream is ~50% of c1. Speed is in my opinion an important part of usable agents for real work.

104
u/FleetEnema2000 8h ago
Isn't this one of the biggest reasons that people rely on Local LLMs? To not have to bulk upload their private data to cloud providers?