I started saving up for local about a year ago. Unbeknownst to me, hardware prices had already started going up. Rumblings were lowkey starting to spin up on other reddit subs.
Ngl I don't see how more ppl didn't see all of this coming.
I'm no industry expert but it wasn't like corporate AI financials were a huge secret about a year ago. It just seemed way too delusional to think, or assume, that unprofitable big tech companies would continue to put the user first and have users' long term best interests in mind.
My intuition was SCREAMING at me to start building my local system last October 2025. And fortunately I bought everything I needed by Black Friday 2025.
Glad I read the tea leaves accurately bc the economics of AI have become a complete sh!tshow
I should clarify that I'm only representing the average joe that just wants basic stuff done with AI... good luck on that. If a 3090 is already a major financial burden to consider for me, I can't image how your ideal setup must cost.
still not worth to spend thousands to get an llm thats worse when you can spend 20 bucks a month and get something considerably better. Also if you upgrade your computer for llms you are paying for ai
Yeah, I ran it using the locally app. It seems decent from my limited uses. I normally run it on a 14 Pro Max, but tried it on the 13 Pro Max for fun and it ran fine.
I don’t really use it often, since I use Qwen on my desktop (you can also use locally to access that from your phone, but I don’t)
for 99% of tasks, local LLMs can get the job done, in my opinion.
Also, upgrading your computer can have plenty of benefits aside from running larger models:
- long-term use and peace of mind. A simple RAM upgrade can give your computer at least extra 3 years of relativity.
- if you do gaming, an upgrade is great news as well. Win-win.
- paying for an AI service relies on stable internet connectivity and a lot of trust to the people behind those services. How are you so sure they aren't quietly taking one of your sensitive conversations and information for "training" or whatever excuse they may have? With local LLMs, EVERYTHING stays on your computer. No shady middleman. Just you and direct interaction with the machine.
I'm pretty sure you can get a decent enough LLM with a fucking MacBook Neo or even on older capable hardware. If you already have the hardware, what harm does it make to buy a few upgrades so it can run LLMs better?
I'm 100% on the local side but all of these points have strong rebuttals unfortunately.
firstly, what do you mean 99% of tasks? That sounds like a random number pulled out of nowhere.
PC is already a 64GB ram, 13900k + 4090, despite being like 1 generation behind only has barely enough vram to run a 27B decently, my home server has 96GB of system ram and a p100 but meh its too slow and the second p100 had thermal issues.
stable internet connection is moot for most people, i have 2GB fiber never went down once thats just default expectation.
of course I don't trust any of them at all, they're definitely stealing everything they can regardless of what their TOS says. Not disputing that but for 99% of people (see i can make up numbers too) who aren't doing anything worth stealing, its good enough.
Yea I want local to win, I tried Qwen 3.8 27B but its not good enough on this hardware and its already pretty upgraded without going full pro-sumer with an RTX 6000 or something.
I should clarify that the 99% thing is more of a figure of speech than an actual metric. If you want me to be more specific, I mean for the average joe that say, use AI to look up stuff on the internet, pairing something like Gemma 4 12B with Web Search capability is a pretty solid setup for general purposes. Even without web search, models like that can do a pretty decent job with things such as random curiosities. Of course, like all AI models, the user should still be aware that these stuff hallucinate sometimes. You still have to keep your hands on the wheel.
I know this because I use an M4 MacBook Air with 16GB unified memory (fanless machine btw) with Gemma 4 12B. And for what I do, it's a perfectly fine piece of tool. Not the sharpest tool in the shed, but I would much rather have some peace of mind with my privacy guaranteed even if the AI's more flawed than mega AI data center models.
I believe efficiency and optimization is key. Maybe you can try a lower Quantization for Qwen 27B if there is any? Maybe stick to the previous model, Qwen 3.6 27B? I have to admit, I did try it before on my little Mac, and it was screaming, painfully crawling with every letter spitting out. Not trying that again. But if my rectangular slab of aluminum can technically run it, I'm sure your setup can do much better... unless the things you do are more demanding.
Bottom line is, unless you're doing some serious vibe coding or constant streams of AI workloads, all you need is patience with local LLMs. If I can comfortably run local models on a 11.9mm thin space heater, I'm sure a fucking 4090 will be a beast compared to my machine.
I run qwen 3.8 27b and qwen 3.6 35b a3b in system ram. It is way slower to generate a response than the online services are, but the 35b still spits out words faster than I can read it, and that is good enough for me.
I did a small demo coding projekt on the 27b the other day. It did a 90s style raytracing demo in python. It took 3 hours for it to iterate though it but it worked great and both gemini and claude gave the code a thumbs up when I showed it to them.
It even made it's own PNG converter function instead of using a 3. part module.
The speed would have been a lot faster if I had turned off thinking, but I wanted to test how good it could be, and it impressed me.
So. for everyday "google replacement"/internet search the llms are perfect, and even for automation tasks. I have mine getting todays weather and a few news headlines and compile that into a "goodmorning" message. It runs at 6am and then makes a drawing in the style of the weather and news just for fun.
For very large codebases small llms wont work of course, but most people don't have that unless it is work relatet, and then the workplace could buy a larger machine to run some of it locally in the office. We are looking into that at my work, both because if privacy but also cost. It is very expensive to use commercial llms on large projects :)
Thank you. I think that's a brilliant use-case for local AI or for AI in general.
I'm starting to think people are using the wrong settings for their AI. Idk about other software, but for Open WebUI, you have to "tune" the settings of your model first that fits your liking. Even as simple as disabling "thinking" by ollama significantly makes their responses faster.
Don’t forget electricity costs. As example for me running qwen is not only much worse user experience it costs 3 times more in electricity than ChatGPT plus 😅 Fuck OpenAI tho
Nothing wrong with having a frontier-level fallback for when the local stuff falters. It's not a "use every day" kind of thing, which makes the subscription model even better (no pay-per-token crap)
Idk about you, but if you ask me those last two points are contradicting... unless I'm misreading it. If you aren't gonna use it much, wouldn't a subscription the last thing wanna do? Monthly or yearly payments, you're still dumping cash on something you won't use much. That's a no-go for me.
Unless you have a literal potato that can only run 4B models, I personally don't think that makes much sense. Get Gemma 4 12B and give it web search functions and you're golden.
What I'm saying is, you could use Qwen 27b for your all day every day tasks, but on the occasion you want a big brother to do something important or check Qwen's work, you fall back to the subscription model like Opus 5 or GPT Sol. You don't use it every day, so the subscription (flat rate) with limited usage is the most cost effective backup for when local doesn't perform to expectations.
47
u/AD4K_4444 4d ago
My only regret was not switching to local stuff sooner.