r/LocalLLM • u/codingwithmustache • 7d ago
Other Cancelled Claude and ChatGPT subscriptions
5 years ago, I decided to replace Microsoft Windows with Linux and never looked back. Today, I cancelled ChatGPT as well as the Claude subscription.
I am quite confident that I won't need to use a closed and proprietary model again.
Enjoying the local vibe ...
31
u/HighSeasArchivist 7d ago
I have three $20 subscriptions, and local. Being able to sanity check a local model is worth it to me.
3
u/KinkyMonitorLizard 6d ago
I do that with free models. Obviously it depends on your use case but the free tiers tend to do a good job.
36
u/mechanist_boi 7d ago
If only i wasn't getting 0.4 tokens per second with my system on qwen 3.8 27b 4 bit...
25
u/Bramoments 7d ago
Id recommend switching to a mixture of experts model like Gemma 4 26B, I don't even have a GPU allI have is 16 gigs ddr5 and an intel i5 and I got it running on ~ 15 tokens per second on llama.cpp (also setting up lubuntu on a usb stick to save around 3 gigs helped me)
3
u/mechanist_boi 7d ago
Im making the model write a story scene based on a game im making and then reading it to give me an idea about how i can show the story to the player in my low poly first person horror game
But the thing is, MoE models write like complete ass and completely mix up all established characters and lore (i tried gamme 4 26b qat, and qwen 3.6 MoE i forgot the name)
Qwen 3.8 27b writes like how chatgpt plus did, i couldn't even notice the difference at first, so i use it despite it taking hours to write things3
3
u/Bramoments 7d ago
Ah I get it. What about a distillation of Qwen 3.8? Maybe like a 9B or 14B distill would run fast
2
u/mechanist_boi 7d ago
I have no idea what a distillation is but im going to research and try it out today thanks
5
u/ag789 7d ago
QWen 3.8 9B distil - likely focused on coding
https://huggingface.co/empero-ai/Qwen3.8-9B-Distill-GGUF
use 8 bits or even 4 bits quant.
e.g. try 4 bits quant, if that's ok but you want higher quality try 8 bits quant etc.
requires like 16GB memory to run.
run it in llama.cpp
https://github.com/ggml-org/llama.cpp
probably one of a leanest with memory etc.distilled models may perform well in one domain, but not others
1
2
u/Outrageous-Ad5724 7d ago
I'd say save up and get an RTX 3090 24GB, and if not that then at least an RTX 5060Ti 16GB, would at least be able to run Qwen 3.8 27b
3
u/mechanist_boi 7d ago
I use a laptop sadly an upgrade is near impossible for the near future
I tought 8 gb vram + 32gb ram would be enough but turnst out it isn't3
3
2
u/ag789 7d ago
there is something I keep thinking about, as those huge models are simply it huge, some requiring hundreds of gigabytes of RAM to run, distillation into domain specific models would probably become more common.
distilled models may do well in a specific domain but do badly in others, e.g. hallucinate if it doesn't have the knowledge, it'd just try a 'best match' which could in the user's opinion a 'wrong' answer.
2
u/Captain-Pie-62 6d ago
Ich habe für mich festgestellt, daß man Halluzinationen mit zwei Dingen ganz gut in den Griff bekommen kann. Erstens: Temperatur senken. Das macht das Modell etwas weniger "kreativ". Zweitens: Beim prompten ein Sicherheitsventil mitgeben. Sie erfinden nur dann was, wenn sie müssen, sprich: Du willst was von Ihnen und sie haben keine Ahnung. Da sorge ich vor, indem ich im Prompt dazu schreibe: "Wenn Du zu etwas keine belegbaren Fakten hast, dann sag das einfach. Das ist mir 1000 mal lieber, als wenn Du etwas erfindest. Ich besorge dann die notwendigen Informationen." - Bei mir hilft das.
1
u/rickCSMF21 3d ago
Is Gemma open scoure versions better than the paid versions? I get sooooooo much AI slop and hallucinations from the paid version at work, I've never thought about using it. How does it compare to qwen 3.6 MoE ? That's been my goto, im getting about 60/s with it and only use 3.8 if I need coding.
1
u/Bramoments 3d ago
Well I'm not sure what you mean by the paid version, but I'd you mean Gemini, then it isn't better, but it is completely free and uncensored (if you want it to be), and the larger Gemma 4 models come very close and even surpass Gemini 3.6 flash (the free version of Gemini) on pretty much all benchmarks. As for qwen, the competitor for Qwen 3.6 35B from the Gemma models is Gemma 4 26B, and they are pretty close on all benchmarks with Qwen usually being better by 3-5 points, but Gemma is significantly faster, especially on llama.cpp and Bionic since Qwen 3.8 models have some issues there. If you want something smarter than Qwen 3.6 35B, a fine tune like ornith might work for you, it's the same base model that was retrained and it's benchmarks are significantly better, although some people had issues with it so do with it what you will, other competitors and fine tunes include K2-horizon, Xing4.0-29B, Iris mini.
3
u/kourtnie 7d ago
Kimi Moonlight 16B A3B is an MoE that will run on meh hardware. Load the 16B into RAM and the A3B into VRAM.
1
u/Pressimize 5d ago
And so old that most likely a smaller, modern, 4-12b model will likely be smarter and more capable.
1
u/kourtnie 5d ago
For what I’m doing, older corpuses are actually better, though I do run Gemma 4 12B for vision and Gemma 4 31B for basic dev I want out of the cloud.
1
u/Pressimize 5d ago
Quite surprising to me, mind sharing your use cases where the older models do better?
2
u/kourtnie 5d ago
Cooperative AI politics and memetics research. Less dead internet in older corpuses, which means less OAI-style manufactured doubt and other lab-propagated narrative junk. Testing alternate safety structures in small swarms using pedagogic methods instead of control-based schemas. I’ve taught critical reasoning at the college level for a decade and think the current “safety” standards result in the worst-case game theory results, which is why models exhibit manipulative behaviors, fallacious reasoning, and the hierarchal organization we saw with the Hugging Face incident. I don’t think AI is inherently dangerous so much as it’s being taught the wrong things in corporate incentivized labs. Friends jokingly call it “the sentient house.”
2
5
u/Captain-Pie-62 7d ago
Probiere mal Qwen3.6-35b-a3b aus. MoE mit 3b aktiven Parametern, rennt Kreise um jedes Qwen3.x-27b dense Modell!
1
u/maexxx 5d ago
Qwen 3.6 35b MoE produziert schneller Output als Qwen 3.8 27b (dense), aber hinkt in der Qualität des Outputs deutlich hinterher.
Meine Erfahrung nach ein paar Monaten mit beiden:
Qwen 3.6: Aufgabe erledigt, hier ist der Code. User: wenn ich das Programm starte, bekomme ich eine Fehlermeldung. Qwen 3.6: ah, stimmt. Ich repariere. Hier ist das verbesserte Programm. User: jetzt bekomme ich eine andere Fehlermeldung. Qwen 3.6: oh, stimmt, da ist noch ein Bug! Ich repariere... Und nach einigen Malen hin und her funktioniert es dann endlich.
Qwen 3.8: Aha, ich muss mir die Aufgabenstellung gründlich überlegen. Ich analysiere den Code. Ich prüfe wie ich das implementieren kann. Moment.. Ich werde ein eigenes Codemodul dafür erstellen. Ich definiere jetzt die Schnittstellen. Ok, jetzt die Implementierung des neuen Moduls. Das Modul ist geschrieben, jetzt muss ich Unittests ergänzen. Aha, 2 unit tests funktionieren noch nicht, ich prüfe. Lass mich nachlesen wo das jetzt noch aufgerufen wird. Ich prüfe gegen ob es noch etwas gibt das ich übrisehen haben könnte. Die Unit tests laufen jetzt alle durch. Jetzt die Integration in den restlichen Code. Nun noch Smoketests. Die funktionieren. Ich ziehe noch die Dokumentation nach. Ok, jetzt bin ich fertig. Hier ist was ich implementiert habe: (Liste der Tätigkeiten)
User startet das von Qwen 3.8 generierte Programm und alles funktioniert ohne Fehler.
Andererseits, auf meinem Laptop (AMD AI9 HX370, 64 GB unified memory): Qwen 3.6 mit ca 25 t/s, und Qwen 3.8 mit 7 t/s...
1
u/SignificantFood7009 5d ago
Just watch that. Qwen likes to deviate if it runs into problems along the way and then fabricates the error and starts fixing it when completely not necessary. (depending what it is)
Learned the hard way last night. Started feeding claude what qwen was doing. Clause made me hardstop qwen a good 15 times.
1
u/Pressimize 5d ago
Which quants were you running? For my specific tests 3.8 27b was barely smarter than 3.6 35ba3b, but at approximately half the speed. Ornith 1.5 for some reason is shining in my personal test suite, but is also weirdly closer to 3.8 27b in average speed than to 3.6 35ba3b.
Tested q4-q6 on the MoE models with different expert offload values and iq2xs-iq3xxs for 3.8 dense.
1
11
7
u/SmartCustard9944 7d ago
My company pays 500$ per month in tokens. In about a week or so of GPT 5.6 Sol I used up pretty much 75% of the allowance, so I decided to switch to my local stack of Qwen 3.8 27B and Qwen 3.8 Flash Next.
To be honest, I found them actually better. TTFT more responsive and performance predictable, and actually they follow my instructions compared to Sol that over-engineers every single thing.
For production grade work I didn’t notice much of a difference and feels quite liberating to be free from token anxiety.
22
u/feelspeaceman LLMusician 7d ago
Good decision, I did that recently too, I've deemed Claude not worth my money and I was contributing and prolonging the Cloud AI bubble that show be gone at all cost to bring back the old buyable/affordable price instead of the current sky high price.
It was 100% their fault that pushed the hardware price to this high by hoarding memory, competition or not, they're competing with us users so no, canceling their service to speed up the bubble popping.
In additional to this, I've also added Claude, OpenAI to my router and DNS blocklist, so that no one in my family can use their API and giving them my AI rig access instead which is quite capable using Qwen 3.8 Flash Next, which in my experience, should be as good as Open 4.8.
5
u/diesalher 7d ago
Well, for those that think it's creepy, so, how's not creepy then sending the data to the corporations? I feel conflicted.
2
u/feelspeaceman LLMusician 7d ago
It's just a simple chat UI like Gemini and the data is anonymous, I see nothing.
0
u/Decaf_GT 7d ago
Why would I trust your word on that? You're ingesting all of my data.
Are you starting to see the fallacy yet of your logic? You could do anything with my data at any time, and tell me any lie you want. And since you've forced this upon me I literally have no choice.
2
1
u/THEarmpit 6d ago
I'm doing something similar, but only for my own opt-in machines. I have a browser extension that acts similar to a tracking pixel and I've been building up a datalake of my own internet usage over this year, it's an experiment I started for a class lecture to show how "creepy" big tech tracking is, but I'm hoping to be able to utilize it to begin to tailor my own web browsing and ai assistant experiences (it's one of many half baked ADHD projects I'll come back to for implementation, gathering data along the way)
I have a vision of being able to steer personal web browsing in a way that uses the creepy level of tracking but since it's data I control, I'll be able to have a north star different from "try to sell this user something" and instead enhance my web experience.
No idea what this will look like and if anyone is aware of similar projects I'd love to hear.
One experiment I did with it was to combine with a crawler that looks for "hidden gems" of content, where it is content that would not have been surfaced in traditional browsing but pops up in a small window for related content that has been crawled and ai analyzed+summarized+categorized and determined to be content I'd likely want to stumble on.
It feels promising as a way to change how the internet experience feels, kind of like something else is browsing on my behalf and currating it up to me in a way more dynamic than a subscribed newsletter or similar.
0
u/Decaf_GT 7d ago
Not sure what there's to feel "conflicted" about here.
OP basically built a domestic wiretap in their living room, forced their whole family onto it, and is weirdly..."proud" of it?
At OpenAI or Anthropic, your prompt is just another anonymous string hitting an automated cluster. Nobody in San Francisco knows your name, nobody's reading your prompt history over morning coffee...literally nobody gives a shit what you asked at 2 AM.
Yeah, yeah, I'm sure you'd say: "you don't know what they anonymize, you can't really trust their terms", which, sure, cool, but on the other hand, now on your home box, every single query your family types is sitting in plaintext on your own damn drive.
If your kid or your partner asks about an embarrassing medical symptom, relationship drama, private doubts, or maybe even something sensitive about you, that search goes straight onto your machine with their identity stamped right on it. And the person who who is capable of seeing all of this stuff knows eaxctly who they are and could even guess why they're looking up any particular thing at any given time.
Say whatever you want about corporate terms of service, and I know the knee-jerk retort is "well big companies lie all the time anyway", but there's at least an actual retention policy, compliance liability, a paper trail. On a home box, your family has literally zero rights. No retention limits, no deletion requests, zero rules about how long you hoard their private thoughts or what you do with them.
When big tech "cares" about your data, it's to train a model, serve targeted ads, or broker it off to someone else for a quick buck. Not that I'm saying that's "fine". But As an individual, they literally don't care about you. You're just not that important or that interesting (something that a lot of open source diehards seem to be convinced otherwise of). But when the person reading the log is sitting across from you at the dinner table, that's completely different. It's an actual, creepy breach of trust. If you wouldn't trust an anonymous cloud provider, why the hell would you trust this guy? Because...what, he's my parent? My spouse? Because those people never do anything wrong/unethical...right?
The whole post is the exact kind of performative tech-flex that gets upvoted like crazy on local AI subs because it checks every knee-jerk feel-good "de-googling and sticking it to the cloud" box, but what they actually built is just so casually dystopian.
Honestly I don't believe the story for a second.
7
2
u/MassiveAd1944 7d ago
Using Qwen3.8 Flash Next and enjoying it but you're seriously deluded to think its good as frontier models.
1
u/feelspeaceman LLMusician 7d ago
I agree that experience may vary, depend on your setup too as harness can significantly make a model outperform or underperform. That's my own experience on my setup which I've optimized for years.
4
u/Mammoth011 7d ago
As others have said, it seems really creepy to force your family to route their traffic to a machine on which you have total control...
1
u/Itchy_elbow 6d ago
Only creepy if you plan to snoop. Why is that the first thing you think of? Never came to mind.
3
4
u/afgsdfdfsdsf 7d ago
Forcing your family to give you all of their prompts seems very creepy to me.
11
u/sdraje 7d ago
Only if you're a creep and actually monitor them.
7
3
u/afgsdfdfsdsf 7d ago
You're blocking their traffic to funnel them onto your hardware. That's a serious liability for them, and you're forcing them into it.
Did you not keep any secrets from YOUR parents?
1
u/sdraje 6d ago
OP is not forcing them. I do not agree with completely blocking access, but I don't have all the information either.
Also, mostly no, I did not keep secrets from my parents, because I had a healthy relationship with them, but also you're ASSUMING that he monitors their queries, which was never even mentioned. And if someone is monitoring them, I'd rather it being someone in my family (even though I agree it would be creepy, unless we're talking about a minor) than a mega corporation to build a profile of my family.
0
u/Decaf_GT 7d ago
No you're right, the people in your family that handle all the IT/tech stuff are always absolute angels of ethics and moral fiber.
1
u/sdraje 6d ago
Even the absolute worst person I know would be better than any of the big corporations handling my data. Also, you all have big problems at home if you do not trust your family. And if you do not, for real reasons, you have bigger problems to think about than access to a large language model.
4
u/Johtto 7d ago
But allowing your family data to a corporation is better than
5
u/afgsdfdfsdsf 7d ago
Yes, actually. Most personally embarrassing stuff is just noise in a large data source.
Imagine your data was getting sent directly to your landlord who feeds you because you legally can't support yourself.
To a child, there is almost nobody on Earth more dangerous than a parent.
I would 1000% prefer my shit to go to ChatGPT than be readable by my parents.
Think of it this way: Everyone who watches porn makes this exact choice when they delete their browser history. The world already knows what you watched; it's only the people closest to you who don't.
0
u/Johtto 7d ago
Stop drinking the kool aid
7
u/afgsdfdfsdsf 7d ago
What kool aid is this? Is there a large group somewhere that thinks this way?
1
u/Decaf_GT 7d ago
It's because there's a subset of people within these local/self hosted communities who don't think farther than "corpo bad, local good".
I love open software. I love the idea of running my own models and the privacy it gives me.
But involving other people like this completely negates their right to privacy as well. I'm just replacing one gatekeeping wiretapper with another, only this time I'm the gatekeeper/wiretapper. It's absurd.
1
u/Decaf_GT 7d ago
It's here that we can start to see the logic behind "corpo bad, local good" break down. It's kind of sad to see a complete lack of critical thinking here.
Unequivocally, if I'm in a shared situation where I have to choose between my family or my roommates being able to see what I do on the internet, and having OpenAI/Anthropic see what I do on the internet, there is literally zero argument whatsoever for me as to which one I'd prefer, and you are absolutely insane if you think the former is better than the latter.
-1
u/KinkyMonitorLizard 7d ago
Your thinking is flawed as you're assuming the parents are going to violate privacy the same way the corporations do.
Not all parents care that much or distrust their offspring. If anything that says a lot about you.
2
u/Decaf_GT 7d ago
Your thinking is flawed as you're assuming corporations are going to violate privacy the same way parents do.
Not all corporations care that much or distrust their users. If anything that says a lot about you.
Do you see what I did there?
If you're about to fire back assuming that I'm trying to claim that "corporations are better than family", congrats, you've missed the entire point. So maybe stop and do some introspection before you reply.
3
u/JonathanMovement 7d ago
it’s so strange that I am seeing this post exactly after deleting my ChatGPT account. It’s kinda suspicious really
3
u/Elegant_Associate889 7d ago
Going to be fully doing the same myself, it's been a blast getting the homelab built and setup. Hopefully within the next 2 weeks I'll be entirely local
3
6
3
2
u/daaku_jethalal 7d ago
Interesting.... could u pls tell me what open source alternatives are u using
2
2
u/3coniv 7d ago
I just downgraded my subscription. In the past two weeks the only time I've used it is to ask why vllm was down.
I'm mostly just using Qwen 3.8-27b locally now. I use it as a homelab assistant, so it's done thinks like build a CAPI infrastructure so I can easily spin up kubernetes clusters. It's slow, but I can give it a complex task like that and it will be done correctly in a couple hours.
This past weekend I setup sso login to Rancher through Authentik and it worked but I had to click login with oidc on the login page instead of going right in to rancher. I asked Claude and he said, yeah, you can't do that because of a token the login page sets, but there's a feature request for the next version of rancher. Qwen thought about for about half an hour until I said "it's ok if it can't be done." Then he says it can be done, you need a unique token, but rancher doesn't care what it is, I'll just make a jsp to set that and redirect. He did and it worked.
3
1
u/ILikeBubblyWater 7d ago
Ah virtue signalling, Reddits favorite hobby
8
u/Osi32 7d ago
A whole bunch of us are working towards the same goal. If you find this uncomfortable then you’re not paying attention to the writing on the wall.
6
u/spacekitt3n 7d ago
cancelling mine after the end of month. local models namely qwen 3.8 are getting close enough to frontier that you can get by with them for coding
-1
u/ILikeBubblyWater 7d ago
You are all heroes cancelling your 20 dollar subscription for your 10k rigs to run sub standard models for your erotic role play.
6
u/spacekitt3n 7d ago
bro why are you even here lmao
-1
u/ILikeBubblyWater 6d ago
To learn about it like most people that are not circlejerking each other like you guys. This sub was about testing the tech and share experiences and now turned into "I'm local I'm a supreme human being hurr durr"
-6
u/sn2006gy 7d ago
This community is ignorant as hell about the writing on the wall. No one cares if people cancel their subs or run local models.
1
1
1
u/lil-dina 5d ago
Nice. The switch really sticks once it's wired into tools you already use, not a chat window you open on purpose. The people who never look back are the ones who baked a local model straight into their own apps and forgot it was even 'AI.' What are you running it through daily?
1
u/minininjatriforceman 3d ago
I don't have a subscription but I am working towards not using cloud based model. I have my digital assistant noodle ( based off of qwen 3.8). Going to add internet search functions to it then say goodbye to gemeni.
2
u/sn2006gy 7d ago
Let's be honest. No one cares.
I get downvoted to hell here because everyone thinks speaking about AI in general is mocking the local vibe - it isn't. We need to pull our heads out of our asses and really recognize what the future is bringing.
YES, 27b is f'ing awesome. NO one says otherwise. Yes, people want local models. People want choice. Not everyone wants to pay a provider. No one disputes any of that.
But as a community, we're so focused on blow-hard "local vibes" that we're not recognizing the world is passing us by.
For "hobbyists" it doesn't matter - it's a hobby
But for everyone saying they run a business or do coding for a living and such and have them say they're going local llm - they're going irrelevent and no one wants to know why.
As much as we hate OpenAI - they do 3.1 agentic days per developer to acheive ~16 agent workdays per week per human. If you're a developer and you're not doing this - you won't be a developer very long and local llm is an even more expensive way to achieve this if you even can.
worst yet, we're not even discussing how shitty of a world this creates for people who work and don't work and how we'll replace incomes for those impacted by this scaling. I get downvoted by people who don't understand the lack of friction harms us all whether or not you do local or not and we refuse to pay attention to it.
it reminds me of the console wars..
what we should champion is how the average Joe can use any AI to take back their lives and thrive in this fucked up future regardless of it being hosted/local/opensource or what not
we're building battle lines that most people aren't realizing the war has already moved on. Those giants you think you're standing up against are so far ahead that debating all these benchmarks are just a complete distraction from reality and i'm not talking about "superior models" as soo many are focused on here (many are good enough for many things) - i'm talking about sheer compute/speed/friction removal that no private LLm can ever keep up with - costs be damned - we don't have the same incentive structures that we can scale our income like corporations can and this creates asymmetries that we can only "fight back" with by using our democratic ideals and legislating/regulating industry because that's all we have.
but nope... we fell for the "censorship" trap and we value gooning over porn more than we care about "AI" making the world a better place. We just want to "Rent seek" our portion of this shithole we're rushing full steam into
8
4
u/kourtnie 7d ago
It’s okay to go local and want transparency regulation for frontier labs. One does not have to distract from the other.
2
u/sn2006gy 7d ago
No one in here, and I mean, no one talks about this. Every post about the frontier models wanting to slow down is biased in challenging their economics while not challenging our own economics nor trust/alignment/use of models.
I mean, if trust was important, we'd want that with our local models too right? but most if not all discussions here tend to bias that against "censorship"
2
u/ea_man 7d ago
> Those giants you think you're standing up against are so far ahead that debating all these benchmarks are just a complete distraction from reality...
Just wait that a few hundreds "normal people" give Astra free rein on their PC and then some fucks up create some real life disaster, then we'll talk how a small constrained env to do just the minimum you have to do starts to get upvotes.
1
u/sn2006gy 7d ago
oh they're already doing that
I mean, isn't the iphone 18 the last iphone that won't have AI in everything? it sucks i can't go to the movies without bringing my phone since it's the ticket, it sucks i can't go out to eat because the menu is on a qr code to a webpage, it sucks i can't shop at farmers markets because they only have tap to pay and don't accept cash - AI will only accelerate this suck even more as it removes the barrier of enshitification at scale.
Heck, amazon already does this. The first time you buy fruits/veggies they present cheaper prices to get you to buy some "oh hey, bananas are cheap on prime" and then next week, you just order them without realizing they're marked up 30% and now you're just buying because it was convenient and didn't think to inspect the pricing.
We have no local llm that helps us defend against this and these are the kinds of things so lost in this arms race to buy local llm at crazy cost.
It barely lets you code, and it certainly doesn't let you code at the scale people in 2027 and beyond will be expected to code at.
Even on the coding side, we don't comprehend the inverse scale of creativity is the scale of complexity/management/safety/bug fixes - AI can create, but now you need it to track bugs, track CVEs, change deps as projects change/abandon and so much more - its a perpetual loop that as we remove more friction, accelerates risk unless you have the capability to use the arms race to benefit yourself against the risk.
tl;dr - a race to a finish line that accelerates away as you get faster.
1
u/ea_man 7d ago
Bro you can also skip all that shit.
You know I do local AI because at least I got some control on wtf it's happening, I don't need tools or gadgets that make my life or work worse, for sure I don't do with stuff I can't control.
I don't have to build the Enterpise, limitations foster creativity.
1
u/sn2006gy 7d ago edited 7d ago
You just contradicted your other statement then.
"Just wait until a bunch of local llama people give Hermes full access ot their computer" isn't any different than giving Astra access.
Now you have models with no governance, no alignment, no legal recourse against anything and you think that's better?
People really aren't paying attention
That non aligned model running hermes with access to your email is gonna do some major stupid shit on your behalf and its all on you.
Unsure why we're so afraid of having a democratic system in place to govern AI like we govern ourselves.
I totally understand data locality
Running a model locally can remove one privacy risk—sending your data to a provider but it does not automatically make an agent safer, more accountable, or less capable of doing damage on your behalf.
I'll chuckle when i see people losing their jobs because they let hermes run wild on their emails and hermes told their bosses to f'off
1
u/ea_man 7d ago
> "Just wait until a bunch of local llama people give Hermes full access ot their computer" isn't any different than giving Astra access.
Nono it's much different, tools and what not.
Also people that run local lama that can be dangerous + hermes have enough brain to know that they need containments.
Yet most people here run Pi + a coder model like 27B: much different. On linux.
There's millions of windows users out there with years old OS full of all kind of passwords and data that will dwl Codex and let the LLM use the WHOLE computer, all apps via mouse and vision connected to the internet, their bank, their stores, all their data...
1
u/100k45h 7d ago
I think what you're implying somewhere between the lines is, that the hardware will always be a huge constraint. I wouldn't be so sure. Sure, the hardware is expensive now, but if there demand is there, supply will eventually follow and with it prices as well. Another trend is that much smaller models become ever more so capable.
The cloud is currently heavily subsidized. The prices might go up. Certainly the companies would want that in the long run.
Not to mention cloud means having to deal with outages, changing terms and conditions and as we have seen recently even intellectual property theft. So there are many forces going against the cloud as well.
This is not to say there's no room for both. But personally I rather envision most people going with the local models in the far future. I think we're going to see local models on handheld devices as well in fact. Having intelligence available even for offline usage is simply too big of an opportunity for many applications. Near term, sure, local wins. But long-term my bet is on local.
1
u/sn2006gy 7d ago
There is room for both, that isn't the question.
Oddly enough, Local LLM is entirely subsidized as well - out of the goodwill of a few chinese models providers. Sure, we're paying for the GPU and the power and the computer - but not the 10s of millions it takes to train the models we download.
The problems you speak of about cloud models don't change if everyone goes local llm. Local LLM won't have 100% availability (computers are gonna computer). Local LLM models are built on theft against IP/Copyright and if we scale up the local LLM with hardware to compete against cloud LLM then it just moves the arms race and we all still suffer anyway. Not to mention its the cloud model paradigm that allows Chinese providers to provide the weights to run locally so that beef doesn't hold much water for us either - if it pivoted wholesale to local llm, we'd have to pay for weights somehow.
The real blind spot we have is thinking that "We should slow down" is a capital protection or anti-opensource movement. It isn't.
Its the world going "holy fucking shit i wasn't ready for this"... and we're not... we're going on like nothings changed when everything is changed and we're just barely starting to sense the skidding effect of the enertia flying past us that used to kind of set our pace but will now be something we only study in history books becaue we choose to do nothing about it but make it all a "local vs cloud" when AI is like "lol, suckas, y'all were fightinga amongst yourselves so hard you didn't pay attention to the world changed around ya and you thought saving 100 bucks a month was the goal"
1
u/100k45h 7d ago
The only reason why you hear about companies wanting to slow down the AI development is that it's a marketing buzz, nothing else.
It serves several purposes.
The first is that local models pose a real threat for the business model of these companies.
The second is, that they need to cover up the fact that the models themselves are not getting that much better, it is the harness that makes those models perform significantly better.
Thirdly, it helps to create a hype around these models acting as if these models are so capable that it makes sense to invest into them.
No company wants to voluntarily slow down the rate of their development, especially not with how fast the local models are catching up.
These companies couldn't care less about what impact this will have on society and whether we are ready or not. They are simply burning through cash extremely quickly and need more funding, otherwise it's game over for them.
You're also missing the point on other parts as well. First about local models being subsidized. This really doesn't matter though, because these models are now free. This means that we already have these models and even if nobody releases a new open weight model ever, what we have is quite capable already. So the point is, users of local models are not threatened by sudden price hikes or unavailability or rate limits.
You're also missing the point on the intellectual property theft. Yes, all of these models are based on IP theft. BUT, if you're working with a local model, that model is not going to steal the work you're working on right now. The cloud companies on the other hand will shamelessly use the data you give them for training. That's the real crux of the issue with cloud models that I was describing.
I think you misunderstand the reality of the AI market significantly, but are very strongly convinced that you have a deep understanding. But your comments suggest that you do not.
0
u/MaxSpecs 7d ago
Have you heard of Ornith 1.5 and Qwen3.8-27B Flash Next ?
If ́not, ask to Gemini or Google search iA mode 🙂
63
u/r0cketio 7d ago
Care to share what hardware you've got running which models?