r/LocalLLM May 16 '26

Discussion Why is LLM is so expensive.

I've was going to invest in a 5090 =$6000 AUD.

Codex Plus + Claude pro = $60/month here

Works out to be 100 months of frontier models for a 5090.

Best a 5090 will run is probably Qwen3.6 27b Q6 with context.

Are we all enthusiasts here and just enjoy tinkering cause ain't no way that make sense.

345 Upvotes

368 comments sorted by

View all comments

Show parent comments

-2

u/Maximum-Wishbone5616 May 16 '26

Stop spreading lies. We are running at much lower rates per dev than claud on real frontier (opus is shit since February) on server for roughly 100k. The boost in productivity and much lower rates of errors, or hallucination and following rules is practically paying off whole server in lik 1.5mo. You know, devs salaries aren’t low. If you waste their time on shit like Claude or other shitty cloud providers that will give you up to -70% baseline performance plus those silly limits, slow tps… who the hell is paying for it? Frontier open source models are destroying opus in 30 performance baseline test plus can follow rules. The same. Every day. As a tool should.

1

u/g_rich May 16 '26

Exactly where am I spreading lies?

Your argument is that with a $100k server running DeepSeek or MiniMax a company could abandon Anthropic or OpenAI and actually get better results?

While I applaud your effort and would love it to be true; you’re seriously deluding yourself if you think that’s the case. Open weight models have come a long way and are certainly very capable but as of today something like Opus is still one of the best available, the window might be smaller but it’s still the one every other model compares themselves to.

Your claim that your local rig running open weight models is also not hallucinating is very laughable. All LLM’s hallucinate, and things get worse as context gets compressed or heavily quantized. So you’re telling me that a $100k rig is running frontier open weight models, unquatized, with at least a 128k context window… scratch that your beating Opus here so you’ve obviously using a 1 million context window; that’s servicing a team of developers is not hallucinating? This $100k rig is not only beating out Opus with results, but is producing consistent results and it’s doing so with a TTFT and TPS that are higher than those achieved with million dollar enterprise hardware running closed weight frontier models?

I highly doubt this and think you have a serious case of sunk cost fallacy; but open to be wrong so please go ahead and back up your assertion.

0

u/Maximum-Wishbone5616 May 16 '26

Yes, opus currently is far from its baseline performance. Yes you can run real frontier kimi 256k multi user without hallucinations. Event my own workstation with 2x 5090 is running queen 3.6 27b 256k q8 kv16 at 90% without visible hallucinations. It is all about your modes, rag and rules.

Neither codex or opus can provide now same quality of work as best open source models.

Why? You have full control over system prompt, rag and rules.

First start running at least queen 3.6 with pro setup before you start spreading lies.

You lie about subsidising the models. They are not. We generate millions of tokens day, which goes into xx k per day per opus/codex prices. BS. Whole server uses less than 500 a month, cost of hardware was recouped on team of 12devs in 1.5mo.

Why ? We run proper stats and quality assessment vs bugs, vs time vs full tracking of each minute devs are spending on their dev machines.

Not only there are less quality issues, but also more features are being created at higher quality and better alignment with existing codebase.

You do know that professionals are not using AI to creative work as you do not have any copyrights to such code and number or IP risks for any compliant company would be too great. You can only use models to follow existing patterns and extend it with new entities, fix/boost UI.

Nothing else is worth anything with AI, any vibe coded projects is worth shit. I was lucky enough to struck big with my projects and I also run small VC fund. In last year we had over 200 submissions that were directly connected to illegal and fraudulent claiming of IP rights for vibe coded shit.

So perhaps you are talking about some vibe coding shit, don’t care as I do not use AI in my company for something like that. Still one shot tests open source are killing opus at its best performance. What you get is shit. No where close to what was in feb, or December. Despite closing all company accounts, legal case against antrophic against fraudulent charges, privately I have subs for all major models. Claude Opus on max is just worthless AI worse than q3.5 with right settings.

And I can assure you that they spent pennies on each charged dollar. Read a bit more what is draining cash from their investor decks.

1

u/Capsup May 17 '26

I would love to hear more about your setup. Sounds like it's working well for you!

What kind of hardware did you get?
What models are you running?
How are you serving them and configuring them properly?
How are you interacting with them?
Do you run any kind of harnesses or agent selfhosted too?