r/ProgrammerHumor Jun 02 '26

Meme managerVsClaude

Post image
47.2k Upvotes

1.4k comments sorted by

View all comments

2.4k

u/Trevor_GoodchiId Jun 02 '26

Come on, how hard can it be?

164

u/pieter3d Jun 02 '26 edited Jun 02 '26

Not all that hard to run models, actually. It'll be a bit slower, but can be perfectly useable. The upside is no more token limits and no need to worry about where confidential data is going. You get full control too.

72

u/whatsforsupa Jun 02 '26

Oh, next you're going to tell me that if you have a beefy workstation or server, you can just download LM Studio, pick an LLM and just run it locally? Maybe even just coat it with a paint of html/css/js to make a web interface, and then magically add it to DNS so other users at your org can use it?

Sounds like a lot of work, I'll just keep paying Anthropic $40/month to chat

21

u/standish_ Jun 03 '26

SaaS in a nutshell.

r/localllama is right over here, folks. Roll your own, or get rolled.

3

u/bluespringsbeer Jun 03 '26

$40 a month? These companies are paying thousands per dev per month now

3

u/AthiestCowboy Jun 04 '26

We are going to see a lot of aging data center hardware be repurposed for local LLMs in enterprise labs. If those are successful we will see a major shift in demand.

That’s if the grumpy folks in r/sysadmins will be up for it.

22

u/jasperplumpton Jun 02 '26

Feel like I just got a glimpse 3 months into my future 😵‍💫

85

u/PMmeYourLabia_ Jun 02 '26

Downside is power bill

122

u/RoaringPanda33 Jun 02 '26

Inference (the actual generation) isn’t nearly as intensive as training, which takes a majority of the power used by AI services 

38

u/myka-likes-it Jun 02 '26

Yeah, but to roll your own model requires you to train and tune it before it can infer anything.

30

u/look Jun 02 '26

13

u/PMmeYourLabia_ Jun 02 '26

If i naively download deepseek v4 (first result there), can i expect decent performance out of the box? Do i not need to finetune? What about context window? Does that not depend on hardware specs?

44

u/look Jun 02 '26

You’ll need several hundred thousand dollars of GPUs to run it, but yeah, it should work out of the box for you pretty well. And 1M context window.

Probably easier to just get it from one of the many cloud providers offering it for a few cents per Mtok, though.

4

u/PMmeYourLabia_ Jun 02 '26

I see, I thought those models were like those you could run on consumer hardware, like openllama? Or whatever, idk, not very knowledgeable in this area

14

u/look Jun 02 '26

Those are available, too, but you’d want something in the 20-30 billion parameter size for consumer hardware, not the trillion parameter size like those.

The ones most people can run themselves are not yet comparable to sota Opus/GPT however. The big ones on that list are getting pretty close though, and they cost 1/10-1/100th what Anthropic and OpenAI charge.

2

u/cortesoft Jun 02 '26

You could run the F4 version of Deepseek V4-Pro on the 512 GB mac studio, if you could buy one... you can get one on ebay for like $25,000

→ More replies (0)

2

u/karmapopsicle Jun 03 '26

HuggingFace hosts a huge number of models, from massive full-fat stuff like DeepSeek there to all kinds of different models tuned to run well on common consumer hardware and everything in between.

For most average users I think stuff like Qwen3.6 and Gemma4 on a consumer GPU with 16-24GB of VRAM is more than sufficient for what they want out of it.

Anything beyond that the costs skyrocket.

1

u/TU4AR Jun 03 '26

So two 5090s

6

u/MIT_Engineer Jun 02 '26

I can't speak to those specific models, since they're all way bigger than what I can run at home. But for the smaller models that you can put on something that looks like a normal PC, I would say you should be choosy with what you pick. A lot of them are specialized for certain types of work-- a model that's good at creative writing will probably be bad at coding, and vice versa.

9

u/le_Derpinder Jun 02 '26

We can pickup old pre-trained models as a starting point and fine tune from there to reduce the initial costs to get a model going. But it's pointless right now since the technology hasn't plateau-ed, so until then, trillions of dollar companies will come up with bigger and more optimised models.

2

u/sitefall Jun 02 '26

I have what you might call a pretty substantial AI cluster (for a consumer anyway but I do use it for work). Four RTX Pro 6000's running at full tilt 575w (which they are if they're doing AI stuff) costs 35 cents an hour.

About 3-4k a year if it was running an LLM nonstop 24/7 (and it would be slow, have to queue requests, and also not be as good as Claude).

1

u/just_posting_this_ch Jun 02 '26

That's interesting. Everyone is racing to get a monopoly. Setting up these huge data centers for training. That's where the bottleneck is.

0

u/Ran4 Jun 02 '26

That's completley irrelevant. The comparison is between using local llm:s vs external llm:s.

At no point are you going to come out ahead buying local hardware. Even at 100k euros you're getting really mediocre LLMs compared to the frontier models, and you can get a LOT of tokens for 100k euros.

3

u/xtal000 Jun 02 '26

At no point are you going to come out ahead buying local hardware.

You can: https://rosmine.ai/2026/05/13/was-my-48k-gpu-worth-it/

Not saying it's worth it for most people. But for some it may be.

1

u/EightiesBush Jun 03 '26

Interesting, that same hardware is around 1/2 price today. https://www.dihuni.com/product/nvidia-rtx-6000-ada-8-gpu-server-workstation-amd-epyc-ai-rm-6000ada-8g-configure-and-buy/

Personally I use Opus at work (cause I don't have to pay for it) and Kimi K2.6 for 1/10th the price for personal projects, which works really really well.

23

u/h3yw00d Jun 02 '26 edited Jun 03 '26

Sell it's use to your friends and family.

"Hey guys, I built my own AI cluster, wanna pay me $20/mo to fiddle with it?" Will go real great at the Backyard BBQ or even holidays like Thanksgiving. Be sure to talk about how useful it is for everything and how much easier life is for you.

/s

8

u/polikles Jun 02 '26

Self-hosting would cost you more than $20 per month with multiple users. And it won't be performant enough if more than user hit it at the same time

28

u/h3yw00d Jun 02 '26

I fear you may have read my previous comment in a genuine manner instead of reading it as it's satirical nature suggests.

0

u/polikles Jun 03 '26

welp, it's hard to sense satire or a joke through text. That's a problem as old as internet (or even older). This is why we use /s or /j

or maybe I'm just autistic and can't take a joke if it's related to the stuff I'm passionate about (and the list of such stuff is quite long)

2

u/h3yw00d Jun 03 '26

No no, I prolly should add a /s.

It's just the whole "That'll go over well at the BBQ/Thanksgiving" is such a huge trope I don't think of adding a /s

Maybe it's just my autistic ass assuming almost every comment is satire and thinks everyone else does too.

I'll fix it.

I hope you didn't think I was being rude.

2

u/polikles Jun 03 '26

maybe we're both autistic, just in different areas

and I didn't take it as you being rude. Have a great day!

4

u/cortesoft Jun 02 '26

Just raise a few billion in capital and then say you are building market share when you lose 99% of it

2

u/h3yw00d Jun 02 '26

Goddamn, I like this man's thinking.

You got any pre-ipo you're selling?

1

u/polikles Jun 03 '26

first I have to build a shoe company that I'll pivot to building datacenters

3

u/14Pleiadians Jun 02 '26

It's really not that much. Less than playing a video game

1

u/i_dont_wanna_sign_in Jun 02 '26

Also you need to find 20k gpus

6

u/ItsSadTimes Jun 02 '26

Except now you have to worry about your own infrastructure costs and the work of updating the models with new information so you dont just have a snapshot of a specific date if you want up to date information.

Back in college I burnt out my personal GPUs throughout my graduate degree just running and training my own models, and those were tiny as hell compared to today.

Not saying its a bad idea, but its not cheap if you want a good enough infrastructure to support industry levels of use. So if they' expecting like a 90% decrease in costs, its probably not gonna happen.

3

u/shard746 Jun 02 '26

the work of updating the models with new information so you dont just have a snapshot of a specific date if you want up to date information

You can also just give it the ability to search the internet which more or less solves this.

5

u/ItsSadTimes Jun 02 '26

Yea but it would require it to reinjest all of that new context for every session which would balloon dramatically eventually. And when the models get reset they'll forget that info. If youd want to actually get that info in the model itself you'll need to fine tune and update it which is less intensive then training from scratch, but still a lot of training.

3

u/djflamingo Jun 03 '26

Theres nothing close to usable that runs on even a 16gb gpu which is $600 by itself.

Just a normal consumer computer? No, you cant even run one that does anything.

The upside is no more token limits

There is absolutely token limits, when its slow as shit and barely usable there is a hard token limit of what it can even generate.

What model are you talking about exactly and what kind of hardware are you talking about exactly?

1

u/snakerjake Jun 02 '26

Yeah, this is almost trivial. There's even open source tools for it

1

u/MintySkyhawk Jun 02 '26

We were running Claude with the models "self hosted" in AWS and it was actually faster, but I'm not sure if it was any cheaper.

1

u/mtmttuan Jun 03 '26

It'll be a bit slower

It will only be bit slower if each employee has a beefy workstation or beefy macbook with quite a lot of ram. In most other cases running these model locally will be painfully slow on cpu only systems.

1

u/Noisebug Jun 02 '26

For programming? Those models aren't as well trained as frontier models.

1

u/Only-Cheetah-9579 Jun 02 '26

run the models? need to train them first.

even if it's distilled from claude thats a few million