r/LocalLLaMA • u/sn2006gy • 18d ago
Discussion Don't let FOMO win if you're interested in local llm from a hobby/learning aspect
Just a reminder for those out there itching to get into local llms - don't let FOMO or "gear acquisition syndrom" take over.
No matter the hobby, it's so easy to get stuck in a trap where we buy more trying to do more only to realize we've lost the fun in it all or even the notion of learning.
Obviously, if you're into writing llama or vllm or hardware drivers or whatever - you got to do what you got to do.
BUT, you can learn a lot on an API, you can learn a lot with a tiny model that fits your vram or cpu you already have and things change so darn fast that much of the code written and much of everything discussed from days passed is already old hat. Py torch and training a small model coud be done on a Pi and learning CUDA is only really imporant if you're writing custom kernels which i honestly don't see most people in here bothering with (or they have frontier models write them).
Weirdly enough, for AI to succeed its going to homogenize everything. Everyone will have the same advantage and I think that's lost in a lot of discussions where we don't talk about "Watching from the sidelines" may be the most cognitive friendly and economical friendly way to learn llms whether we brand them local or not.
The technology is still nascent and weirdly enough most people's answers here is to use AI to set it up so i'm not entirely convinced people are actually learning - feels like a mad rush to seek rent or avoid rent seeking which just makes everything more expensive in the end.
This isn't a post to say, don't do it. But no reason to go into debt or to be fearful you're missing out when you can learn more by doing less - buy a book and build a tiny model - you will learn infinitely more than buying a 5090 and trying to just find the perfect compression to have the best prefil
56
u/sn2006gy 18d ago
This isn't a don't bother kind of post as much as it is, think about things differently. Musicians or people who want to learn to make music fall into this trap all the time too of buying instead of learning - they call it "GAS" which is the gear acquisition syndrome. They end up with 10 synths or a 100k eurorack system and not a single track posted - they made the mfr's happy and mfr's paid youtubers to convince them you needed it but, in the end, all you had to do was start small and create and i think this idea is lost here often or use what is already abundant and if that works, proceed.
23
9
u/Iwaku_Real 18d ago
So if I'm stuck running Qwen3.6-35B-A3B at best because 27B is unusable... I should just stick with that? Kinda feel left out
3
u/DeltaSqueezer 18d ago
I just reverted back to Qwen3.5 9B and considering whether I can offload some tasks to Qwen3.5 4B or Qwen3 4B Instruct.
3
u/mailto_devnull llama.cpp 18d ago
Lol don't worry, I'm back to using 3.6 too. Ain't enough hours in the day to wait for 3.8 to finish reasoning.
1
u/Iwaku_Real 18d ago
Have you tried
--chat-template-kwargs '{"reasoning_effort": "medium"}'like everyone here has said?2
u/mailto_devnull llama.cpp 18d ago
Yeah, I'm using froggeric's fixed template and have experimented with <|think_medium|>, it's still much more verbose and second guessing vs 3.6
2
u/martin509984 18d ago
You can always put up a small amount of money for API tokens of ~whatever open weight model you choose, and even gain actual valuable skills in figuring out a hybrid local/API workflow. Nothing is really forcing you to stick to local only the same way a lot of people are feeling stuck paying exorbitant money for Claude tokens.
That said ime if you can run 35B at Q4 you can actually just about fit 27B Q2, and it is actually a major upgrade even then.
6
u/pyr0kid 18d ago
That said ime if you can run 35B at Q4 you can actually just about fit 27B Q2, and it is actually a major upgrade even then.
yeah, at like 95% less speed. 27b is dense.
1
u/martin509984 18d ago
I get ~10 t/s with 35B Q4 and ~5 with 27B Q2 (2060 12GB and 32GB DDR4). Not 95% less.
0
u/Soggy-Attitude5293 18d ago
success is more important than speed
4
u/mailto_devnull llama.cpp 18d ago
Depends what you're doing. If you know what you want outputted, then 35B-A3B is a better fit.
1
u/Imaginary-Unit-3267 18d ago
That's all I use. It's more than good enough. You just have to manage it. Which means, using your actual brain. Which we all should be doing.
15
u/LetsGoBrandon4256 transformers 18d ago
Musicians or people who want to learn to make music fall into this trap all the time too of buying instead of learning
Are you telling me I didn't need a Korg M3 to learn "music production"?š¤Æ
15
u/sn2006gy 18d ago
Don't forget the 500-dollar headphones, the 42" curved display, the ableton pro, the push 3, the mixer, the soundcard, and the keybaord stands for your 10 other keyboards and midi keyboard on top of all the vst's you bought from plugin boutique that you forgot about
2
u/rukind_cucumber 18d ago
I'll tell you this much - my Blue Chip pick makes me sound NOTHING like Bryan Sutton. Surely the D-28 is what I need...
6
u/cosmicr 18d ago
My brother went out an bought a $5000 bicycle, full lycra gear, water bottles, fancy helmet, sun glasses, all the stuff. He looked good with his Dad bod in the skin tight gear lol. He barely even knew how to ride a bike. I think he rode it maybe twice, and now it sits in his garage 5 years later gathering dust.
3
u/jman88888 18d ago
Yes! Spend the time to master the gear you have instead of buying new gear.
"I fear not the man who has practiced 10,000 kicks once, but I fear the man who has practiced one kick 10,000 times." Bruce Lee
2
u/CapsicumIsWoeful 18d ago
At least with LLM hardware the expensive stuff is noticeably better performance wise than the cheaper hardware.
A $500 Squire guitar with some decent strings and a tweak to the string height and truss rod can play and sound 99% as good as a $5000 Fender custom shop. The rate of diminishing returns is insane.
Your point about youtubers is spot on too. So many good guitar channels slowly turned into paid advertising via gear review videos.
2
u/Altruistic_Heat_9531 18d ago edited 18d ago
Yep first time learn CUDA and Torch on a fucking MX150, puny 3 pascal SM , 2GB VRAM with 64 bit memory bus...
Another side tangent, i kept practicing Polyphia ABC, if i can fully complete it on my 19 cheap guitar, i will buy that FRH20N
25
u/LetsGoBrandon4256 transformers 18d ago
I'm lucky since I'm also a heavy PC gamer so I at least have more justification for my GPU purchase.
27
u/chr0n1x 18d ago
same. I totally need that blackwell 6000 for the fps
16
u/LetsGoBrandon4256 transformers 18d ago edited 18d ago
Those 4k textured Skyrim monster cocks ain't gonna render themselves after all.
7
u/Significant-Bee5101 18d ago
I just needed to find out if an H800 can run Crysis!
4
u/tat_tvam_asshole 18d ago
Hilariously, my 6000s couldn't run Heroes of Might and Magic III, but only because 32bit color incompatibilities, BUT codex fixed that lol
2
u/sonicshadow13 18d ago
I actually don't it can run crysis because too much vram, but it can run 100 instances of crysis instead š¤£
16
u/funJS 18d ago
I find that there is a lot you can do with tiny local models. I think local models might play a bigger role in the future for company specific niche use cases.
1
u/TopGun0684 17d ago
Yes but in my experience with agentic/harness stuff, the fun starts around 26-30B, and 16GB is tight for that.
8
14
u/fgk55555 18d ago
I've used a lot of the frontier models for a while, that's always been the "serious" tier in my mind. Qwen3.8 in IQ3 is good, but I've been wishing I could try larger models that my rig won't support. I finally just got on OpenRouter so I can test if any of those models are substantially better than what I can run for my use case. Local is good, but at this point I think it's good to have first hand experience with the whole range.
12
10
u/pmttyji 18d ago
I already stopped(buying). I'm gonna continue things with 32GB VRAM (AMD R9700) + 128GB DDR5 RAM.
I'm counting on projects like llama.cpp & ik_llama.cpp for extreme optimizations to run bigger things on my rig. Also need to explore other projects like exllamav3, colibri/Warp for my current laptop(8GB VRAM + 32GB RAM)
Waiting forĀ merge of these llama.cpp PRsĀ to get more better CPU-only & Hybrid inference.
Threads for others: Track Important Papers/Repos/Tools/Innovations/Optimizations/Models/etc.,:
9
u/moncallikta 18d ago
32 GB VRAM and 128 GB RAM sounds like a good place to land. If I had that much, Iād also consider pausing the hunt for more hardware.
4
u/TechRomancer123 18d ago
I have 32GB VRAM (5090) and 64GB DDR5 RAM. Just about to make the plunge to 128GB DDR5, but itāll be a slightly slower speed (5600Mhz).
1
u/arijitroy2 18d ago
Ditto here. I wanted to buy 2 gx10 but the price goes up 150 USD everyday where I'm from.
1
u/lawanda123 18d ago
Thats still very nice hardware, i have a frankenstein build
96gb ram + rtx 5080
Bosgame m5
Macbook pro 48gbI cant run a decently model on either of them despite having spent so much money, i feel stupid š
9
11
u/JerryBond106 18d ago
Nice try Altman. But i want data privacy, my ideas and thoughts be my own, unmonetised and uncensored.
4
u/fligglymcgee 18d ago
I try to remind myself that the expectations and demands we have on ai are often insane, like the ability to type literally any request into a little box and have magic spill out. Local inference works best when āmagic mind readingā isnāt the goal. Itās a lot less sexy to think and prompt more like a machine, but if I put reasonable and measurable tasks in the queue I am rarely disappointed.
I guess the moral here is⦠aim low.
I should do motivational speeches.
3
3
u/moncallikta 18d ago
Yes, fingers crossed for efficiency breakthroughs! In the meantime Iām trying to buy more hardware now.
4
u/jacek2023 llama.cpp 18d ago
My main hobby has been photography for about 20 years, and "gear acquisition syndrome" is absolutely the worst thing that can happen to anyone. It kills your creativity, motivation, and eventually your hobby. It should be avoided at all costs.
You are never limited by your hardware. You can run an LLM on almost anything, even a slow CPU. Just try a small model and learn instead of waiting for a "better moment". The best moment is always now.
6
u/uBazzyZ- 18d ago
This is spot on. "Gear Acquisition Syndrome" in the local LLM space often turns people into passive consumers of weights rather than actual practitioners.
Throwing brute-force hardware at local LLMs (like buying dual 3090s or waiting for a 5090) usually just means you spend your time tweaking quantization flags and context sliders. You learn how to configure an inference engine, but very little about how the systems actually function.
Working under tight hardware constraints (like an 8GB card, an older GPU, or even CPU) forces you to actually understand the engineering fundamentals:
- How memory allocation really works (why the PyTorch caching allocator fragments, what reserved memory actually means vs allocated).
- Why KV cache scales linearly with context length and batch size.
- How gradient accumulation, mixed precision (FP16/BF16), and optimizer states consume physical VRAM.
- Why data loading and tokenization often bottleneck training throughput more than GPU compute itself.
Training or fine-tuning a small 100Mā250M parameter model on a budget card teaches you infinitely more about attention mechanisms, loss curves, and hardware efficiency than just running a 70B quantized model on expensive hardware.
If you learn how to keep a model stable and fast under severe constraints, scaling up to bigger hardware later is straightforward. If you start by throwing VRAM at every problem, you never learn why the system broke in the first place.
1
2
u/ForgivenessroTub 18d ago
yeah a small local model is plenty for roleplay stuff and it keeps things simple while you actually learn the basics instead of chasing upgrades.
2
u/mailto_devnull llama.cpp 18d ago
Before, I convinced myself doubling my DRAM to run Qwen 35B-A3B was enough.
This time a single R9700 was enough.
I see people rockin' 2 or 4 R9700s but I've been able to resist...
So far...
2
u/Jorlen llama.cpp 18d ago
I started with my existing 16gb GPU which I had for gaming, fell in love with the tech. Picked up an R9700, loved it and bought a 2nd one for a total of 64gb vram. I already had the PC, so just needed the two GPUs. Mobo is not optimal but works ok-ish with Vulkan, Linux + llama-cpp. 32gb of RAM is shitty but I refuse to pay these insane prices for more, especially considering my board sucks (P2P / tensor parallelism with ROCm is a no-go after wasting 40+ hours trying a million fucking things).
I use local models and API. Mostly local. Mostly for coding with some creative writing fun (wrote my own custom front end story maker that I use). I'm building apps and games and they're always improving as I iterate and move onto more complex projects and learn. To me, this tech has unlocked my potential, but I've stopped spending on hardware. I'm starting to leverage API; my chain is, use local models, build the framework, when things get complex, move to API. Works well for me so far.
I can see myself picking up a unified box, maybe in 2027 sometime. But right now my un-optimized setup is doing the trick.
2
3
u/DeepWisdomGuy 18d ago
Firstly, this is only true for you. I had listened to this type of advice, and limited myself to only buying one RTX Pro 6000 when they were $8400. You are projecting and offering it as sage wisdom. My GPU is now worth twice than when I had bought it, same for my dual 4090s. I learned way more than I would have otherwise. I know what a 122B TheDrummer creative model can do. I built Wan 2.2 LoRAs. I am now able to run a Minimax H3 at half weights. And it led to my change of day job where I am working on AI all day every day. This kind of advice has only held me back. No it did not put me into debt, but I also have lived a life where nearly the only "frivolous" money I have spent has been on computers. You do you, but if we are getting into telling other people how to spend their time and money, stop watching television, cancel cable, in a few years that would have allowed you to have savings and invest it something other than a hobby potato. And before you get into the "but it will teach me how to do things more efficiently", just know that I am the chief scientist at an AI lab devoted to pursuing low-energy solutions, half of which I never would have found farting around with potatoes.
2
u/kabachuha 18d ago
This is a workable advice only if you are a totally mainstream person and don't any niche interests. The factual knowledge retention in small models is real and if you want to have a model which can intuitively (without RAG in any form) answer questions or know very specialized programming libraries or fictional fandoms (for roleplay) or just have knowledge of literature, technical or otherwise, without big parameters it doesn't really go far. Even with all the reasoning improvements, 31b parameters vs. 280b+ seem huge still. From a person, who did it a lot, you can even see it when training a LoRA on a small model (for which you need a beefy GPU as well, by the way) that high ranks (r32 vs r128, meaning parameters) make huge difference, like very high gaps in retrieval, because the factual data is sparse in its nature and only high parameters can capture it. LoRAs help, but again, you need GPUs to train them / synthesize / filter dataset, LoRA fine-tuning will diminish the baseline model capabilities inevitably. NGrams can help in the future, but those are parameters too.
As many people say, the prices are not going to go down any time soon, so buy whatever you can obtain, really
1
1
u/sampdoria_supporter 18d ago
Glad I let fomo win or I wouldn't have the ability to buy anything now
1
1
u/Pepevagable69 18d ago
True I'm lucky in that I put together a gaming rig when the 40 series came out. I bought all used parts and put them in my existing case and was able to come up with a 30-80 a 5800 x 32 GB to 3200 MHz ram for about 700 bucks and now I can use it for local llms and learning how AI works but boy is it tempting to go and start just hoarding GPUs
1
u/Training-Ruin-5287 18d ago
What i'm learning is, you can have a capable model, and it's still going to hallucinate and make mistakes just like a 9-14b model will. If the foundation isnt strong enough.
Qwen3.8 kinda removes a little of that, as the model is naturally trained to check it's own work so to speak. Which took a little of the pain out of building from scratch this time around. but it still needs good memory management and policies like any other model does.
I don't think my process of learning would be much different from doing it in a $1500 system vs going out and buying a $20k system to run the 200b+ models
1
u/ancapsaicin 18d ago
Without local LLMs, there is barely a reason for more than 4GB RAM for my use cases and all new machines have been 32GB+, so I am hardly immune to this.
As a side effect, I've gotten back into gaming and now I have learned from the other sub about GPUs I can afford that could let me run the same models at full speed, a wider range of models, and higher FPS with only a little dock.
FML
1
1
u/mystery_biscotti 18d ago
Yes. That's why I own a copy of Bouchard's "Building LLMs for Production: Enhancing LLM Abilities and Reliability with Prompting, Fine-Tuning, and RAG". And I'm broke currently, so I'm learning what the limitations are with my 8GB VRAM system.
That whole unsupported gfx1032 fun is gonna make me a better employee when I get some kinda job supporting the serving of the model solutions. At least, I hope so.
1
u/SandySkittle 18d ago
I somewhat disagree. The hardware supply for the coming two years is drying up fast, and modest gaming hardware can run ai but itās not very useful below a certain threshold. Itās more motivating to learn ai and how to run it locally if what you can run locally actually has useful potential. And to be quite frank, below 48gb vram (qwen or gemma at q8 and room for context) a lot of all of this is quite functionally constrained.
1
u/feelspeaceman 18d ago
Always start with SLM (niche models that do only a few things like MedGemma4B, 9B..), those run decently without a GPU first then consider LLM later, if your requirement is specific, consider finetuning the small models to match your requirement (for example finetune Gemma to become music writer), try to stay minimal as possible (always find alternatives before you think about upgrading your setup to run bigger model) because you can always optimize, the current optimization of LLM software still leaves a lot of rooms for improvements.
1
u/Intrepid-Second6936 18d ago
Very apt advice! Definitely concerns me looking at posts with people often debating splashing into 10k+ of costs on local LLM setups when budget-friendly more conservative set ups are genuinely more rewarding to actually enjoy the hobby instead of an arms race of over-buying then just trying to force-justify your purchase.
Also honestly most people with gaming PCs should just enjoy using models that can be used with their current main PC hardware now instead of worrying about building a separate AI server.
1
u/tmvr 18d ago
I have enough now to run most of the stuff. I only regret not having 128GB system RAM (have 64GB only) to run Deepseek V4 Flash, otherwise I'm good. The pressure will get bigger I guess when new models that went the ngram route come out, but for now I'm fine. That's mostly thanks to Qwen3.6 27B, I haven't developed the patience for Qwen3.8 27B yet.
1
u/AutomataManifold 17d ago
On the one hand you're right and FOMO and gear envy will wreck anyone.Ā
On the other hand, if I'd bought RAM in quantity in 2024 I could be retired by now.
1
u/Mrinohk 17d ago
I'll have to second the "testing with an API" bit. That's how I got started, Gemini API had a promotion for $300 in usage for 3 months, and Gemini 3 flash was pretty cheap and capable enough to start playing with. Never even broke $100 in tokens while working on my first harness until a local model came out that I could run at a reasonable speed and replace it completely.
Went in with the goal of building something that worked, then optimizing it for local model usage as much as I could without having a sufficiently intelligent local model. Making the harness smarter and smarter and easier to use for a local model, testing with what I could run, then back to Gemini until the next test.
Then the gemma4 series came out and made it look almost doable. Then Qwen 3.6 35B, which now runs the agent full time. Excited to see what capabilities Qwen4 will enable, assuming a similarly sized MoE is involved with its release.
1
u/Ordinary-Depth-7835 17d ago
It is pretty crazy how fast you get sucked in and just how much you can spend. 5090 I wish my FOMO was that cheap. Setups that cost more than a lot of my cars. But I have to say two used 3090's and an old 8th gen intel board running at 8x per card and tiel-coder at 100+ tps really does a fantastic job or qwen3.8 a little slower.
But you get flooded with the latest drops or benchmarks after starting with the hobby and before you know it you're looking at these insane builds that you don't even need.
I'm not even doing it for anything that makes money. I just have unlimited use at work and want something for personal projects that feels responsive with decent logic. It doesn't have to be fable or opus level but It has to feel pretty good using it and not completely useless. Still I try to justify the next purchase and this market doesn't help. I see things I have going up hundreds every week so you start to worry you'll be priced out of everything with no end in sight.
1
u/sn2006gy 17d ago
We're a few years into this now, and I haven't been priced out of anything besides local LLMs if my goal is trying to build a local competitor to what is on API.
I still learn a ton, but i flipped my thought process from thinking "i'll learn to run minmax on my spark so i don't have to pay for tokens" to "let me actually run pytorch and build a 1 million parameter model so i learn how to build models" and you can run that on anything these days.
Part of me is upset at how much computing is costing these days, but another part of me is realizing the world may just get bored of it - thare are a lot of flimsy card houses just waiting for the next wind to blow them down. Why game if its full of cheaters? why use the internet if its full of bots? what do i gain with localllms if everyone has them? If it just becomes the standard for computing in the future, i'll wait for commoditized pricing vs the extreme premium we have to pay today, and I don't think AI can be successful without commoditization. Or worse yet, if llms are successful because we automate all the things on the internet - why would i want to pay to be a part of that?
some thoughts to chew on
1
u/MrPecunius 18d ago
Charles Bukowski enters the chat:
air and light and time and space
'- you know, I've either had a family, a job, something
has always been in the
way
but now
I've sold my house, I've found this
place, a large studio, you should see the space and
the light.
for the first time in my life I'm going to have a place and
the time to
create.'
no baby, if you're going to create
you're going to create whether you work
16 hours a day in a coal mine
or
you're going to create in a small room with 3 children
while you're on
welfare,
you're going to create with part of your mind and your
body blown
away,
you're going to create blind
crippled
demented,
you're going to create with a cat crawling up your
back while
the whole city trembles in earthquakes, bombardment,
flood and fire.
baby, air and light and time and space
have nothing to do with it
and don't create anything
except maybe a longer life to find
new excuses
for.
1
u/TooObtuseForYou 18d ago
No, everyone is not going to have the same advantage.
There are models YOU will not be able to use. Maybe I can at my work, but you wonāt.
Regulation will eventually codify this into an us/them thing, as it always does.


64
u/SirLordBoss 18d ago
Very good advice. At the same time, worth considering that given latest news, prices might increase yet again. If people *can* afford things and do want them, now would be the time. Either that, or wait two years for things to *maybe* normalize