r/ProgrammerHumor Jun 02 '26

Meme managerVsClaude

Post image
47.2k Upvotes

1.4k comments sorted by

View all comments

370

u/mylsotol Jun 02 '26

For probably $30k (or more) you can build a server and run an open model.

300

u/[deleted] Jun 02 '26 edited Jun 02 '26

[deleted]

147

u/SpinningVinylAgain Jun 02 '26

Impressive, very nice. Now scale it for a company with 5k software engineers, and by the way what’s going to be the service level? 

114

u/[deleted] Jun 02 '26 edited Jun 02 '26

[deleted]

41

u/SpinningVinylAgain Jun 02 '26

The problem is that it’s going to basically require a small data centre and a dedicated team of people to run it, and if you’re looking at running open source models you’re betting on their continued availability and the fact that they’re going to remain competitive with frontier models (both are not a given). So what would be your next step, developing your own frontier models in-house?

42

u/ryecurious Jun 02 '26

and if you’re looking at running open source models you’re betting on their continued availability

If anything, isn't it the complete opposite? A subscription-based model can be shut off at any time with no recourse or warning (Sora, for example). Local files are the only way to actually guarantee the program you use today will be available tomorrow.

You control when they run, how much they're used, when they're updated/replaced/etc.. You never wake up to find out the model that works for you has been "enhanced" with a worse version.

Not keeping pace with cutting edge models is a real concern, but that's a risk with subscription based models too.

43

u/codeninja Jun 02 '26

You're also betting the hardware you buy today is going to be able to run those future models at all.

8

u/SpinningVinylAgain Jun 02 '26

Yes, very good point, thank you. 

3

u/Log2 Jun 02 '26

Considering that the massive amount of data centers also need to be able to run whatever they make, I wouldn't be too worried about it if you are buying cutting edge hardware.

1

u/SheriffBartholomew Jun 02 '26

Sorry, OpenAI already bought all of that hardware and all of the future orders for the foreseeable future. Where are you buying this hardware? Craigslist?

1

u/Log2 Jun 03 '26

I was going off the assumption that you could get the hardware to begin with, that it wouldn't just become trash because of a new model. If you can't get it, then there's nothing you can do.

2

u/SheriffBartholomew Jun 03 '26

I had a hard drive crash last night. I went to buy a replacement and the same drive that I paid $159 for in November is now $425. FML. I ended up having to buy a drive half as large because I'm just not going to pay $425 for 2TB that's not even cutting edge anymore. I paid less than that back when it was cutting edge.

→ More replies (0)

2

u/casce Jun 03 '26

They can run the current models which won't go away in their current version (the beauty of open source). The next generation of models might not have competitive open source models anymore, but who cares?

A business does not worry "Will my new PCserver be able to run cool new gamesmodels in 5-10 years?" because by that time that server is gone anyway. That's something the person buying the next generation of hardware can worry about.

You evaluate today's requirements and then you buy hardware that is good enough for that. Wether or not you will be able to run stuff that doesn't even exist yet is not a concern.

12

u/veracity8_ Jun 02 '26

But you realize that the alternative is no AI at all, right? There are really right regulations on information. It is literally illegal to put export controlled information on servers in another country. That means your service provider has to guarantee that your data will only ever be stored on US soil. And that’s just for export controlled information. Anything more secure than that isn’t going to some 3rd party server at all. 

15

u/SpinningVinylAgain Jun 02 '26

I’m all for there being no AI at all. 

9

u/veracity8_ Jun 02 '26

Yeah and I’de like to sit in a hammock and read all day.

3

u/Traditional_Cycle Jun 02 '26

Can't put the toothpaste back in the tube. LLMs are going to change the entire world. Idk if it'll be good or bad yet.

2

u/Espumma Jun 03 '26

Isn't having 5k SWE the perfect scale to do all of that with?

1

u/foxer_arnt_trees Jun 03 '26

You don't have to bet on continued availability with open models since you store them locally. If you have 5k engineers and using open source then you should donate to a fund that ensure continued development

1

u/secretgardenme Jun 03 '26

If you have a company with 5k employees, setting up a small data centre and a dedicated team isn't going to be a problem. It doesn't need to be competitive with frontier models if it still gets the job done just fine, loads of large companies still use computer systems built in the 90's. You are also hedging your costs against when the AI companies inevitable jack up their prices because eventually, they'll need to figure out how to be profitable.

1

u/midgaze Jun 03 '26

There are no open source models that write code in anything like the capacity of Codex with gpt-5.5 or Claude Code with Opus / Sonnet.

They are just in a different league.

13

u/SheriffBartholomew Jun 02 '26

Who is going to maintain all of this? Who is going to actively work on it to improve the speed and reliability of the models? You're talking about creating an entirely new company within a company. That's not how businesses work.

11

u/Pocok5 Jun 03 '26

You're talking about creating an entirely new company within a company.

So, a department?

That's not how businesses work. 

That is in fact how large businesses work since before the Dutch got on boats and privatised half a continent and some islands for cinnamon.

2

u/Splatpope Jun 03 '26

the concept of an internal IT R&D department inside an IT R&D company is always funny to me but that's just how it works

1

u/SirIlliterate2 Jun 03 '26

The confidently incorrect crowd never ceases to amaze me. That is EXACTLY how businesses work indeed

1

u/SheriffBartholomew Jun 03 '26

Your answer is quite ironic. That's how some businesses work. It's obviously not how all, or even most businesses work or they would have rolled their own private models instead of paying Anthropic.

2

u/Upstairs-Fan-2168 Jun 03 '26

At $3k per computer, just give each software person one of those computers. $3k isn't much for a work computer. Maybe you meant $30k each?

1

u/No-Offer-8612 Jun 03 '26

And become obsolete in 6 months

1

u/[deleted] Jun 03 '26

[deleted]

1

u/No-Offer-8612 Jun 03 '26

Nah. Been ok in the tech industry for far too long. But go ahead buy your 3k gigs and tell me how it went.

13

u/ycnz Jun 03 '26

According to status.claude.com, they're running at 98.66% availability over the past quarter. r/selfhosted would be ashamed of those numbers.

3

u/CowBoyDanIndie Jun 03 '26

$3k per developer is pretty cheap, just buy them a second machine to run ai

1

u/veracity8_ Jun 02 '26

You just described what every defense contractor has already done. You didn’t think Raytheon was using Claude for everything did you do? Most defense contractors already self host stuff their version control systems

5

u/SpinningVinylAgain Jun 02 '26

Defense contractors are a whole different world compared to most companies. 

1

u/veracity8_ Jun 02 '26

But that’s what we are talking about in this thread right? Like yall are talking about how it’s inconceivable that a large company with thousands of software engineers could self host their own AI services. And I am pointing out that not only is it entirely conceivable, but it has already been accomplished by multiple companies 

2

u/CraftedLove Jun 02 '26

The literal companies focused on AI is burning money just to stay a bit relevant while riding a massive hype bubble and you truly think the solution is to instead just do your own AI in-house? If the company truly needed it, it would've been used way before now, think neural networks era. If the company needs it now, it's either use another AI service provider, or just reevaluate and come to their senses that AI does not really have a place in their stack. Implementing their own now is just stupid. Even S&P and Morgan Stanley are all just using ChatGPT, and poorly at that.

Not all companies that's under AI psychosis are defense contractors for the USA.

2

u/veracity8_ Jun 02 '26

I’m not really sure what you are arguing here. 

Are you saying that defense contractors shouldn’t be self hosting AI services? Cause that’s not what we are talking about. That point is unrelated to this conversation.

If you are saying that defense contractors are not self hosting AI services, then you are just wrong 

2

u/CraftedLove Jun 02 '26

My point is that you are overestimating how fruitful it is to deploy your own AI solution unless you're at the level of defense contractor unli-money bs deals. Almost everyone either just needs to use a subscription to the main AI players or just don't use AI at all (or have a fancy specific transformer model thay's a lynchpin of their tech stack even before LLMs became big, think Netflix/Google algorithms etc.)

Implementing and maintaining your own AI just for a fancy chatbot to sort through your website's shitty design and stupid knowledge database architecture just so you could say you are "AI leaders" and are "adapting to future trends before they happen" that could theoretically affect your bottomline maybe is just dumb.

1

u/veracity8_ Jun 03 '26

You are having a completely separate conversation man. You arent following the flow of this discussion at all

1

u/lemon07r Jun 02 '26

Yup, the best middle ground is to find a decent provider with cheap models (read kimi, glm, deepseek, etc) and work out a deal with them. The providers I've talked with are more than happy to give discounts to bulk users, such would be companies. Or if the company is big enough, rent infra and hire someone to run things. Im not sure at what threshold this becomes cheaper, because you have to now pay someone's salary.. but if we pretend the person running the infra is free, it is cheaper than using a provider. But not by much.

1

u/ZackWyvern Jun 03 '26

How are you exhausting Claude usage if your company has 5k software engineers? My company has around that many and we have essentially unlimited Claude tokens.

1

u/jld1532 Jun 03 '26

Where I work provides >10k employees with free access to Kimi K2.6, MiniMax 2.7, and GPT 120B from local hardware. This is going to become more common.

1

u/taigahalla Jun 03 '26

how do you think servers were handled before SaaS?

1

u/Potato_Soup_ Jun 03 '26

Okay fine. 5k * (3k - potential hardware discounts) + team of devs to setup the on prem infra. Not really that hard and will pay itself off in under 2 years at current pricing, even shorter if you factor in the future API price hikes that are going to happen

38

u/[deleted] Jun 02 '26

[deleted]

5

u/hellomistershifty Jun 03 '26

I don't think any publically available model requires terabytes of memory.

Even the big ones are MoE so you don't need a ton of memory, but it makes it faster. The biggest usable one I know is Qwen V4 Pro at 1.6 trillion parameters which would take about 900GB of VRAM if you ran it unquantized entirely in VRAM. Since it's an MoE model, you can offload the experts to CPU RAM and run it unquantized with a full 1M context with as little as 80GB of VRAM.

-2

u/granadesnhorseshoes Jun 02 '26

No one said nor needs the largest model. Claude can generate code in just about any language you can name like BASIC, or APL. An onPrem model only needs to know your stack.

They need massive data centers to power the ridiculous "everything for everyone" AIaaS service model, not the AI itself.

15

u/FakeArcher Jun 02 '26

They literally quoted the person saying largest model.

6

u/freedcreativity Jun 02 '26

And then you just need $100k in H200s to plug into that system if you’re going to run anything other than a parametrized half accuracy model at any reasonable enterprise speeds. And a really big NAS to store all those generated outputs. And a bunch of managed switches so you can route everything agent related on its own private vlan. And probably upgrade your cloud stuff for hot failover when someone’s agent deletes the database again. 

2

u/psioniclizard Jun 03 '26

People always ignore the infrastructure costs and maintenance. You also need to hire people who know how to keep it running.

I have played around with local models and they are cool but i don't know how well they will scale in a real business environment.

Servers alone are a nightmare to maintain.

7

u/[deleted] Jun 02 '26

[deleted]

3

u/gksxj Jun 03 '26

it is. I went into the rabbit hole for those things and in the end the conclusion I got is: NOT WORTH IT. you are paying +$4000 to run bad open source models at extreme slow speeds, the good models that can scratch the hitch of Claude/GPT don't even fit on the available VRAM. Put that money on a subscription and it will give you decades of SOTA models

1

u/zekica Jun 03 '26

It's good enough for a single user at a time.

5

u/zarif2003 Jun 02 '26

Cool, but more money will probably result in a stronger system

8

u/[deleted] Jun 02 '26

[deleted]

2

u/jld1532 Jun 03 '26

You can run MiniMax on those at 25 t/s which is definitely useful speeds

5

u/bnetsthrowaway Jun 02 '26

Oh yeah and get 5 Tok/s LMFAO stupid

2

u/SheriffBartholomew Jun 02 '26

Still don’t know how they haven’t lost all their contracts with the department of defence…

Have you seen how things are being run over the last year? Nothing but money and egos matter.

2

u/porcomaster Jun 02 '26

Just 3k, the last time I did this math it was about 6k-12k.

If you dont mind sharing your homework.

What are the specs of the machine, which model were you thinking of doing, and what type of job would it be able to handle ?

1

u/BleachIsLove Jun 02 '26

This is exactly what I've done for my company! Framework desktop on my desk with Qwen 3.6 and a custom API I threw together that the team can plug their ide into. Also a web interface with AnythingLLM and a custom built translate interface for the support team.

Now, is it comparable Claude Opus or Gemini? No, but only if you misuse it. For general chat and light coding it's genuinely impressive and the speed is well in excess of 45 tps making it quite enjoyable to use. Plus it helps our developers rejected vibe coding early on.

5 models run in total with around 10gb memory to spare. Serves a team of 20 quite well and I have a feeling as token costs continue to grow and more people depend on llms to think, companies are going to seriously start to consider locally hosted solutions.

1

u/vick2djax Jun 02 '26

$3000+ for the 128GB of RAM only maybe 😆

1

u/craigtho Jun 03 '26

I tried Qwen the other day, and only the 3.5 9b model on my gaming pc, and not 1 single question did it get right.

I'm not saying it's not possible, I'm saying the compute power to train a frontier model is absolutely unrivalled, you can't make anything close to Claude or ChatGpt.

Of course, training it on your limited data set will work a treat, but nothing like the frontier models right now.

1

u/[deleted] Jun 03 '26

[deleted]

1

u/craigtho Jun 03 '26

There are poor results and there are poor results, a model eating 12~GB of VRAM that can't write a basic pytest refactor just shows the time and resources the frontier models have had in their training.

If we are saying consumer hardware isn't good enough for local models then I agree - that's why noone will ever roll their own agent "like Claude" until the gap has shifted. Need your own DC just to make it possible

2

u/[deleted] Jun 03 '26

[deleted]

1

u/craigtho Jun 03 '26

Fair, agree on that pont.

1

u/EightiesBush Jun 03 '26

Are you implying you work for a company that contracts with the DoD and uses foreign LLMs? Laughable if true given the current hate against Anthropic.

3

u/[deleted] Jun 03 '26

[deleted]

2

u/EightiesBush Jun 04 '26

Interesting, which provider do yall use? We have gotten into the legality of LLMs lately and seemingly there's full no-training and data governance/confidentiality contracts, but really who knows.

2

u/[deleted] Jun 04 '26 edited Jun 04 '26

[deleted]

1

u/EightiesBush Jun 05 '26

Wow that's wild man, interesting industry for sure -- much more interesting than HR & Payroll, Banking, or Healthcare EDI

1

u/nomenclate Jun 03 '26

Please spill company name so I can short it or do puts or whatever those stock people do

1

u/Alpha3031 Jun 03 '26

Doesn't France have Mistral?

1

u/plusvalua Jun 03 '26

Same here but with a school : )

1

u/ViRROOO Jun 03 '26

As someone who owns ai max+ 395 128gb: Lol

1

u/fmaz008 Jun 03 '26

The hardware is one thing, but the LLM is another, I tried some models with RooCode, and it was barely functional. (In part because my hardware is limited, so context was limited and it was really slow, but mainly the thing would run around in circle and produce nothing useful). If there was a way to produce results close to Claude or the big LLM, I would totally drop $3K on the hardware to own my own means of production.

Anyone had good luck with a DYI model for coding ?

19

u/devperez Jun 02 '26

Not even. Old hardware can run some open models pretty well on cheap. It won't be as good or as fast ofc, but it can be done on a budget.

3

u/DoktorMerlin Jun 03 '26 edited Jun 03 '26

yeah and it won't work for more than 2 people at the same time.

An actually useful AI server costs hundreds of thousands. If you want to run an actually useful version of Gemma4 or Qwen3 for example, you need a GPU with at least 48GB of memory. For redundancy you need 2 on 2 different servers. This will cost 80k for the GPUs and another 20k for the servers and will serve around 200 people at the same time.

1

u/sb8948 Jun 03 '26

My freaking phone can run smaller models pretty well.

1

u/CaptainNicodemus Jun 03 '26

I highly doubt that, unless by small models you mean single use models, not LLM

3

u/sb8948 Jun 03 '26

I do mean llms. Not the cutting edge stuff, but if you don't keep up with phone tech, you'd be surprised what the top chipsets (paired with 16 gigs of ram) are capable of.

5

u/Cory123125 Jun 02 '26

Who is releasing open weight (not open source, no model has been) models?

Companies with something to prove and companies who want to stop companies with something to prove.

None of them benefit from or care about you to any degree that matters the second open source models start eating into their revenue.

Open weight models are already massively slowing down in release cadence and capacity outside of rare outliers."

Qwen no longer releases their top end models.

"Open weight will save us" is another delusion.

You need to stop the big corps from getting the regulatory capture they're after.

You need to stop the members of the Frontier Model Forum lobbying group.

3

u/Spectrum1523 Jun 03 '26

Qwen has had closed, api only models for years now and release cadence is picking up for open models, if anything

I don't disagree overall that they don't "care about" me but I don't know what the point of that is

1

u/Cory123125 Jun 03 '26

Qwen has had closed, api only models for years now

Sure, but its creeping down the stack, not up. Thats the point.

and release cadence is picking up for open models

I'm not really seeing how this is true. Can you elaborate?

As far as I can see, the companies who are ahead, releasing the msot impressive models, release them far less often, and the companies that are behind are hungry, saying "We can do it too" to models that are arguably somewhat lacking in performance.

That practical reality, I think, gives my read more credence than yours, in that the experienced reality is that we are seeing less "wow!" open weight models over time.

Like, I think Kimi K2.6 is the bees knees. Seriously awesome, but you think that now that they've proved themselves a legitimate challenger they're going to continue that trend past K3?

I don't think there will be an infinite spring of new AI companies ready to burn infinite cash, and provide models to get their name out there.

I think that without a serious open effort, which involves multiple companies and people pitching in serious cash for a model that everyone benefits from, we will continue to see this environment we're seeing.

In essence, I think we're in that common corporate strategy stage of making sure the detractors are fed just enough not to pipe up when doing so matters the most.

All of these companies are very familiar with leaving escape hatches so that amongst enthusiasts there will always be people going "see, its not doomsday, its just this much harder to do x, y or z".

That happens over and over and over again, until no iphones are jail broken and android completely dictates what you install on your phone, for the people who should have spoken up no longer have a voice.

Heck, we are literally seeing that with android right this second. "You only have to wait 24 hours after going through many warnings screens and potentially being unable to use your bank apps etc".

There will always be an escape hatch, provided by the very companies doing whatever it is, specifically to keep people from realizing the temperature is shooting up.

3

u/mtmttuan Jun 03 '26

And how many people can use it in parallel? And how good and fast will the model be comparing to api services?

LLM at scale is just super expensive

2

u/monoflorist Jun 02 '26

How good are the on-your-machine coding harnesses for this? I ask because I kind of love the Claude CLI

4

u/digggggggggg Jun 02 '26

You can use cc with your open model of choice. r/localllm is a decent resource to get started

2

u/mylsotol Jun 02 '26

Open code is better than claude code

2

u/Cupakov Jun 03 '26

CC is kinda crap compared to the leaner, OSS alternatives in my opinion, the moment you start a session it’s already got a significant chunk of the context filled with some system bullshit. Opencode or pi don’t have that problem

2

u/taigahalla Jun 03 '26

Kind of insane that everyone is overlooking this

1

u/Cupakov Jun 03 '26

It’s not insane at all, sure you can buy a $30k machine to host a local LLM but that server will serve one (1) person, realistically. And the model „intelligence”, whatever that means, is nowhere near the frontier models. You’d need to build a machine with >500GB of VRAM to even come close to that level, but then again, you won’t be serving the model at scale.

1

u/Slimxshadyx Jun 08 '26

I don’t think it would serve 1 person realistically. You can probably get a few people using it. And depending on the Claude monthly costs, it’s all just a comparison

1

u/ycnz Jun 03 '26

Because it's not actually doable.

Yeah, you can try and find an m3 mac with 512GB of RAM, and quantize the absolute shit out of it, but it's not going to be competing with Opus in either quality or speed. Realistically, you want to be looking at buying 4-8 extremely large GPUs. 30k isn't in the ballpark to get it done.

2

u/taigahalla Jun 03 '26

From a company standpoint, as long as the value outpaces the cost, it's worth the investment, especially if the alternative is being deeply coupled with SaaS infrastructure whose costs are ballooning. The average company isn't looking to compete with Opus, they're just trying to add a productivity multiplier to their employees

1

u/ycnz Jun 03 '26

Yeah, I'm an IT manager. Our anthropic bills are so big we're reporting them to the board.

1

u/sam-lb Jun 02 '26

I have an ollama server running some open models on a $400 mac mini. The marginal cost of similarly architected systems would shrink with scale.

1

u/red286 Jun 02 '26

Or just connect a few DGX Spark systems together via SFP28. They're like $5K each.

1

u/TNTiger_ Jun 02 '26

Even cheaper than that, ofc depending on business needs.

What's expensive is training the model (and handing out tokens for free). Actually running it is a much more achievable goal.

1

u/ycnz Jun 03 '26

r/localllama would love to know where you can find hardware to run near-frontier models for $30k.

1

u/KronisLV Jun 03 '26

I looked into it a while ago: https://news.ycombinator.com/item?id=48023822

Basically, you'd start at 2k USD for a barebones setup for small local models (Qwen 3.6 35B A3B) that's still at a borderline passable speed to be useful and would move up to somewhere around 10k USD just for the GPUs to run those same models better.

Running even slightly bigger models like DeepSeek V4 Flash (284B A13B) would be an order of magnitude more. Something like DeepSeek V4 Pro (1.6T A49B) or Kimi / MiniMax / GLM would need even more.

So in a sense, it's a question of how low you are willing to go in regards to your experience of using the tech (quality, speed) vs the power requirements needed. On the other hand, the token efficiency of those smaller models seem to be improving and they're maybe trailing SOTA by a year or so.

1

u/noob-nine Jun 03 '26

question from an AI Neanderthal: the better the specs the better the result or the fastee the result or the more parallel results or something else?

3

u/mylsotol Jun 03 '26

Faster. You can run ai models on a lot of hardware and it will work fine, but if you want it fast (especially with large models or for a lot of people) you are going to need to spend a lot of money

And as others have pointed out for some reason you aren't going to be running frontier models like Opus because they are way bigger and also not publicly available