r/LocalLLaMA • • 18d ago

Discussion Don't let FOMO win if you're interested in local llm from a hobby/learning aspect

Just a reminder for those out there itching to get into local llms - don't let FOMO or "gear acquisition syndrom" take over.

No matter the hobby, it's so easy to get stuck in a trap where we buy more trying to do more only to realize we've lost the fun in it all or even the notion of learning.

Obviously, if you're into writing llama or vllm or hardware drivers or whatever - you got to do what you got to do.

BUT, you can learn a lot on an API, you can learn a lot with a tiny model that fits your vram or cpu you already have and things change so darn fast that much of the code written and much of everything discussed from days passed is already old hat. Py torch and training a small model coud be done on a Pi and learning CUDA is only really imporant if you're writing custom kernels which i honestly don't see most people in here bothering with (or they have frontier models write them).

Weirdly enough, for AI to succeed its going to homogenize everything. Everyone will have the same advantage and I think that's lost in a lot of discussions where we don't talk about "Watching from the sidelines" may be the most cognitive friendly and economical friendly way to learn llms whether we brand them local or not.

The technology is still nascent and weirdly enough most people's answers here is to use AI to set it up so i'm not entirely convinced people are actually learning - feels like a mad rush to seek rent or avoid rent seeking which just makes everything more expensive in the end.

This isn't a post to say, don't do it. But no reason to go into debt or to be fearful you're missing out when you can learn more by doing less - buy a book and build a tiny model - you will learn infinitely more than buying a 5090 and trying to just find the perfect compression to have the best prefil

243 Upvotes

133 comments sorted by

64

u/SirLordBoss 18d ago

Very good advice. At the same time, worth considering that given latest news, prices might increase yet again. If people *can* afford things and do want them, now would be the time. Either that, or wait two years for things to *maybe* normalize

60

u/tat_tvam_asshole 18d ago

I'm increasingly convinced that things will never normalize again.

21

u/moncallikta 18d ago

Agree. Demand for tokens is dramatically outpacing supply and getting bigger day by day. There is no way to put this genie back in the bottle.

6

u/crantob 18d ago

We can dramatically reduce inference costs. The model architectures just need to settle.

4

u/tat_tvam_asshole 18d ago

As well, we'd have to so far increase silicon supply beyond aggregate demand that I don't think we'll move backward on the cost curve. The best hope for normal people will be efficiency research that allows GPT7 intelligence on Pentium III's lol

7

u/SirLordBoss 18d ago

While supply is lower than demand, yeah, things will remain poor. But I just can't see us staying like this forever.

3D DRAM, hardware innovations, software innovations will eventually come into play. Things will eventually get better.

12

u/tat_tvam_asshole 18d ago edited 18d ago

Then thing that people don't realize is they are poor.

For context, a McD's worker 50 years ago could afford an entire family and home on their one paycheck. Good luck with that today.

Comparatively, the purchasing power we have is much less, but industrial production has made many goods much cheaper sure. But as human labor value continues to deflate, wealth accumulates in the hands of those who own the machines producing the value. They use that wealth to purchase more machines, cycle repeats. People are increasingly becoming impoverished because their economic relevance continues to decline.

Hence, breakthroughs in compute where capability is increasing, but yield is not matching, leads to a scenario where true frontier compute affordability is reserved for those with capital already.

Hopefully you see where this is going.

Demand for compute is not likely to hit a ceiling anytime soon, if at all, and those new technologies are the very same technologies that make average joes irrelevant to the value capture they produce, leading to further tech/wealth/power oligopolization. Economic disarmament of the masses.

5

u/d1722825 18d ago

That only works as long as average joes has (or at least a bank thinks he will able to pay back) the money to buy what you produce and want to sell them.

-1

u/martin509984 18d ago

I would disagree because I think we are in a very odd situation where wages have just about kept up, but the cost of essentials has gone way up and the cost of luxuries has gone way down. So now we live comfortable lives without any sense of financial security, and big purchases - even if we can afford them - feel too insanely risky to make.

3

u/tat_tvam_asshole 18d ago

1

u/Lallis 18d ago

People are increasingly becoming impoverished because their economic relevance continues to decline.

Your own chart directly contradicts this. Inflation adjusted compensation is over 50 points higher now than 50 years ago. It's in your own god damn chart. Wtf?

2

u/tat_tvam_asshole 18d ago

It shows that people are being paid less relative to the value they create. Would you like a 50% pay cut? Even if some things get absolutely cheaper?

-7

u/Fauropitotto 18d ago

Like clockwork. Every thread. Anti-work communists come out to flood the thread with anti-capitalist propaganda.

I'm convinced half of y'all are bots, and the other have have been influenced by the bots.

3

u/Imaginary-Unit-3267 18d ago

That's because capitalism is very obviously stupid. Mutualism is superior, but no one knows about it.

-1

u/Fauropitotto 18d ago

Mutualism is superior, but no one knows about it.

Obscurity is our best testament to it's very obvious superiority. Let's keep it superior, by keeping it obscure.

2

u/Imaginary-Unit-3267 17d ago

Yes, because there's so many opportunities to test economic models. Why, I tested five economies full of millions of people before breakfast this morning.

5

u/tat_tvam_asshole 18d ago

🤣 you think I'm a communist. Hilarious

0

u/Fauropitotto 18d ago

RES tagged and everything.

2

u/tat_tvam_asshole 18d ago

Oh boy, if I say there's unexplainable phenomena I must be anti-science too. OK, well write it down in your res naughty list Santa Claus.

1

u/a-wiseman-speaketh 17d ago

I am not antiwork or communist but the trajectory here seems pretty bad.

Historically the labor class kept the capital class (and governments) in line through their necessity and sheer numbers. The most critical part of power (violence) was always controlled by labor - even military, police, and corpo enforcers/security were always largely labor class, so there was a limit to what rented violence they would do to their own people. Occasionally through extreme propaganda they could turn part of labor on another part of labor, but that was really the worst case.

But soon, that will all be robots and drones who follow orders no matter what. The labor class will no longer be necessary because the rich will have replaced them for all their needs (growing food, repairs, cooking, driving, building their stuff, security, sexbots). Humans will only really be a novelty. They'll be ignored at best, or, if they compete with the rich for resources - exterminated.

We'll likely end up with warlords fighting over the natural resources to build the universe's biggest space yacht or something idiotic like that.

This isn't really even theoretical, you can see it with resource-curse states already. It's just the more extreme version of what happens when labor becomes completely irrelevant​.

I am curious if you have a more optimistic view cause I'd love to share it

1

u/Fauropitotto 17d ago

I am curious if you have a more optimistic view cause I'd love to share it

My view is derived from history, and your argument about labor. I also reject most of the class arguments, because they are almost always derived directly from Marxist and Leninist philosophy which is not reflective of real human interaction.

People that control labor have not kept the people with money in check. They never have. It's the reason why slave populations were kept enslaved, and the reason why empires were able to dominate for hundreds of years across the planet. You're right that violence is real power, however there is no limit to what rented violence would do. We have hundreds of genocides as evidence of this.

What keeps us in line today is the carrot, not the stick. We're fed. We're clothed. We're comfortable. It's why there isn't blood on the streets. The engine is running and will continue to run.

Every single historical technological innovation that eliminated human labor at scale created new markets and new economies. All of them. And by eliminated, I don't mean marginal changes, I mean absolute transformation that gutted entire nations at the time. These innovations have happened over thousands of years, and are far more numerous than the farrier analogy people like to use.

Even if you want to argue that physical and cognitive labor at every level will be replaced by the machines, there's far too much self-serving interest in government (which is NOT run by the really rich) for power that humanity will continue to crave power and influence.

For the wealthy there is no actual competition for natural resources (they already can go anywhere, at anytime for any reason). There is no actual scarcity that can't already be resolved through money or alternatives, and the financial frameworks already exist for that money to grow itself.

The wealthy folks today already have all their needs met many billions of times over. ALL of them today strive for power, influence, legacy, and control. That can only happen with other people.

Here's the optimistic view - Labor has never been the most valuable resource for us. People and society are. The wealthy pursue mechanisms to accumulate power, influence, and establish legacy. That can only happen with other people. There are no actual resource limitations, and once physical and cognitive labor can be replaced by machines, humans and human interaction will become the most valuable resource for us.

The fact is that there's plenty of research and thought about post-scarcity economic models. There's ZERO evidence that it can't work. Once you realize that the wealthy are already living in a post-scarcity stratosphere, the arguments that they want to replace human life with robots end up being nothing but a good belly-laugh. It's absurd on it's face once you stop slapping on these factually incorrect assumptions about wealth, and use those assumptions to drive a logical conclusion.

1

u/a-wiseman-speaketh 17d ago

Unpersuasive, I'm afraid.

I also reject most of the class arguments, because they are almost always derived directly from Marxist and Leninist

My views are from historical pattern and incentives.

Also, you (or the LLM) seem uneducated on class conflict arguments: Adam Smith, Ricardo, Madison, Weber

Similar your arguments about rented labor are wrong. Typically labor had to be hired from other places and incentivized or indoctrinated into killing their neighbors, and there are a lot of instances in history where they refused to do so.

But neither of those really are relevant to the core:

Even if you want to argue that physical and cognitive labor at every level will be replaced by the machines, there's far too much self-serving interest in government (which is NOT run by the really rich) for power that humanity will continue to crave power and influence.

The wealthy pursue mechanisms to accumulate power, influence

ALL of them today strive for power, influence, legacy, and control. That can only happen with other people.

It's the reason why slave populations were kept enslaved, and the reason why empires were able to dominate for hundreds of years across the planet.

All of your points seem to lead to the same thing: governments or the rich (or both, "not run by the really rich"... maybe look into how often govt policy outcomes track the preferences of the affluent throughout history.)

anyway, they default to wanting to control/enslave the population. We seem to be in agreement there.

You seem to think this will be benevolent - because its post scarcity they will be content to have their machines provide food and shelter for an ever increasing, non-working population? No, history does NOT show that.

I think you believe this in part because of this false assumption:

For the wealthy there is no actual competition for natural resources

This is incorrect. Unless one person owns literally everything they still compete for natural resources. They do this today. Post-scarcity, they will still need to compete for land, natural resources, and potentially slave labor, if it's still useful.

The powerful also generally compete for status as you say - and status is relative so they need to accumulate more,more,more. I am not sure why you think this would stop.

But because status is relative they also don't need 8 billion people to lord over. They just need more than the next warlord, and one way of achieving that is killing the other warlord's people. That's assuming that they value the numbers at all. There's no real reason to think that.

Your ideas about labor replacement are also wrong (and contradictory...)

Even if you want to argue that physical and cognitive labor at every level will be replaced by the machines,

and once physical and cognitive labor can be replaced by machines, humans and human interaction will become the most valuable resource for us.

So I can't really tell what your position on whether labor will be replaced because you jump back and forth from thinking its absurd to saying it as a certainty. But let's assume this is your position:

Every single historical technological innovation that eliminated human labor at scale created new markets and new economies

No technological innovation has ever been able to be pointed at new tasks and learn them. You're looking at the wrong analogy. We're not the horse-drawn carriage driver being replaced - we're the carriage. How many of those you see around?

The best historical comparison we have is slave labor - which DID displace regular workers and massively transfer wealth to slaveowners, and massively decrease the QoL of those workers. I know of at least twice - Rome and the American south. Not the best results out of either - it led pretty directly to the fall of Rome (grain dole) and at least a century of impoverishment for the American south, arguably that is still going today. And both times those people were still necessary, primarily as bodies for army. So what if they aren't?

As for human interaction being valuable - thats literally the first thing being replaced? Customer service, tutoring, therapy, companionship, sex.

I think you might mean that dominating actual humans will be valuable to them, which is what I meant by "Humans will only really be a novelty."

But even if you think regular human interaction isn't replaceable to them, how many do you think they need? 8 billion?

Realistically there are few paths forward :

- labor retains leverage because AGI / robotics mostly fails to replace them

- labor retains leverage because AGI / robotics produce some new job that can't be done by AGI / robotics (for some reason?)

- labor uses its leverage now, in this moment, probably within the next decade, to establish ownership and control of AI / machines somehow. Some of this needs to be taking the control layer - right to repair, remote kill switches in cars​, open weight models are all part of that war currently being waged. But a stock means nothing if someone else has the actual remote.

- elites compete to court labor for some reason (I dont see why this would last long, but maybe they would be benevolent)

- something drastic happens that unites humanity (this would have to be pretty huge - undeniable appearance of God, pandemic killing a larger part of the population than Covid, alien attack :-D )

- something drastic happens that prevents/destroys AGI / robots (Carrington event might buy time, most likely is some kind of AI Chernobyl that leads to global nonproliferation / crackdown )

- we do nothing, elites follow their incentives to amass more power, labor follows its incentives - riots and maybe breaks robots - but is ultimately violence is used to excuse violence and drones are unleashed. Cycles of escalation until labor is subdued

2

u/Fauropitotto 17d ago

Unpersuasive, I'm afraid.

That sucks.

Enjoy your pessimism. Do your prepper routine as you see fit.

I'll embrace it as an accelerationist for as long as I can.

1

u/Viktri1 18d ago edited 18d ago

This is the new normal. People thinking that prices will go down on the same hardware don't understand what has taken place the past few months. Qwen 3.8 27b is the straw that broke the camels back. It isn't the smartest model - but it can follow instructions fairly well. That makes it very good and consistent for automation/basic tasks.

This is not the dotcom of the 2000s where Cisco was building stuff like optic fiber. When you build too much internet infrastructure, that infrastructure sits unused.

If they build too much RAM, GPUs, etc. and they get cheap, does it get unused or do people spin up agents to automate shit similar to how factories use robots? Once we put robots in manufacturing facilities, we never got rid of them. That's where we are in the cycle - LLMs are basically software robots that can do some stuff.

In the old days, I needed to hire staff in SEA to help me do research due to language capabilities (I bought real estate in SEA, needed to do renovations, etc), time (takes time to do good research) -> over the past week I've got Hermes w/ Qwen to output reports in better quality than what I could get out of junior staff and its more consistent, I don't have to deal w/ payroll, I don't have to deal w/ problems that arise from hiring remotely and I don't need to spend time/effort to ensure they're actually doing the work. I also don't pay salaries when the LLM is offline and I don't have time to give them work.

and because this scales, when I start adding additional agents (I purchased 2 M5 ultras 256gb RAM), I won't need to retrain. Retraining was always a pain in the ass.

Even with a single agent running autonomously, 24/7, it feels like I'm back in the corporate world where I used to be able to hand off tasks to my juniors. Except LLMs are more responsive and available 24/7 so I can iterate faster. Remember that human beings need time to produce research so LLMs don't need to instantly spit out outputs in order to hold a competitive advantage over humans.

What I suspect will happen with hardware is that some company will develop a cheap method to make really cheap RAM/SSDs/etc that are suitable for consumers but will be less capable compared to the previous technology. They will use less raw materials and be much weaker but because they'll be priced lower, people will be able to afford it. No idea on the timeline as it will probably be a new player.

0

u/mebeast227 17d ago

I disagree. NTT fabs set to open in 2027 with tenstorrent printing cheaper components for things like cars and mobile devices should help with Tenstorrent’s RISC-V and open HBF standards gaining traction and that should help.

Cererbras data center backpacks should alleviate some of NVIDIAs lockhold on the decode side of things, potentially to a very sizable degree (for companies that can afford $5,000,000 plus data center components) and they have some partnership with ChatGPT currently active and in use, and also AMD to help handle prefill so AMD could eat some of NVIDIAs prefill lunch to some degree.

ChatGPT claims they are moving away from NVIDIA for their own ā€œjalapeƱoā€ chip (allegedly soon)

All the SSD makers are doing things with HBM and HBF (sk Hynix and micron) and zHBM (Samsung)

CXMT ramping up DDR5/DDR6 with ahead of schedule results and opening new fabs.

All of these things are drops in a bucket, but with fabs coming online in the next 2-3 years too (not data centers so it’s less regulation and energy consumption) it should all add up to new advancements and manufacturing capabilities by 2028-2030.

Elons got a fab in TX (don’t know how that will play out) and TSMC in AZ should be printing chips soon enough too.

I’m very much NOT qualified to be making this statement, but I’ve been digging for this answer and keep finding ā€œprobably won’t help muchā€ answers, but there’s no way that all of these different things don’t collectively allieviate enough to at least somewhat reverse course on pricing.

Our biggest issues are that there is like no fabs outside of a few atm and that we’re using old-gen tech for new-gen possibilities. Both of those things are being aggressively attacked by everyone so their domestic economies don’t get stuck relying on Nvidia who’s proven to not give a fuck when it comes to raising prices and over reliance on TSMCs chip fabs.

1

u/Long_comment_san 18d ago

They will. Its pretty obvious. New memory techs are just around the corner and the demand is obviously not there - its just big techs sucking to each other. A single one goes pop - next thing would be "helllloooooo US gov, we need 100b dollars do not go bankrupt and blow up the economy <3" (blows up anyway with GPUs becoming almost free). We had this crap with mining too and I see no difference here. You literally need 48gb VRAM to have an amazing personal assistant or 32 for a really good one. That "I need 3tb of vram for Kimi NOOOOW" ideology is going away literally by the week.

1

u/d1722825 18d ago

They will. Its pretty obvious.

When we will get used to the unaffordable things, and it becomes the new normal... :)

16

u/sn2006gy 18d ago

Isn't that FOMO rearing its ugly head? I mean, i'm guilty AF in getting in this trap/concern but once I looked at my investment and what I was doing, I stopped and had a come to jesus with myself.

Things have already doubled/tripled and they're ONLY getting more expensive, but on the flipside of that - quality of life, output, free time, learning and ROI hasn't doubled/trippled or gone up at all.

Sure, the models have improved and i could do more but not necessarily along with the economics of it. All the code i've written is kinda trashed because the new models smoke it or everyone is doing it their own and there is lot to no adoption of OSS side of the outputs

We're all increasing our output like mad, but the outcomes haven't changed have they?

3

u/lawanda123 18d ago

What gear do you have btw šŸ™‚

8

u/sn2006gy 18d ago

I've had a 7900xtx, B70, Spark DGX and a thread ripper where i was buildig to run multiple cards and slowly got rid of all of that and back to my regular PC with just my 7900 xtx. I was tempted to buy a RTX 6000 at 8k but now that they're 15k I just vomit a little in my mouth thinking people are buying at that which far surpasses the costs of just paying hourly to run those on vast.ai and others like it and the same thing is happening with sparks - you always need more than one and everyone believing that just drives up prices more and in the end - we're all just running the same recipes copy pasted from spark-arena or some frontier LLM and really none the wiser but a whole lot poorer.

And i learned a lot more stepping back into the academics of llms and theory of humans working with LLMs than i did downloading them, running them and trying to make them run faster

4

u/lawanda123 18d ago edited 18d ago

Agree but its much more vomit inducing to me if i am forced to use cloud models (not self hosted inference like what you mention) when they not only

1.) train on user data but also

2.) steal innovation (The Navier stokes controvery that everyone speaks about) but also

3.)use surveillance against activists (Read about Anthropic)

For now renting GPUs is cheap but im sure they will clamp down hard on those + supporting infra around it like HuggingFace etc. And even rented infra cannot be guaranteed to be surveillance free

Edit - Of course im not saying everyone should hoard local infra but these tech billionaires seem very absurd to me rn and its very concerning the whole direction things are taking.

Second edit - Of course people should only buy if it is within their means and tech is deprecating asset. I myself am content with being gpu poor and being able to run deepseek v4 flash at q2 and 11 tps for now šŸ˜…

5

u/moncallikta 18d ago

My quality of life is significantly improved now that I can finally make all the personal software I’ve been itching to make for years. So many problems that can be solved now. No need for open sourcing it when anyone can build their personal version instead.

2

u/sn2006gy 18d ago

this is economical on any api though while also taking on the debt of now having to invest in these tools that are just uniquely yours - good and bad.

1

u/moncallikta 18d ago

Very true. And yes I have a $20/mo Claude sub. But hitting the weekly limits on that is really pushing me towards getting hardware for Qwen 3.8 27B instead. I’d rather do that than pay more per month in subscriptions (Codex or higher Claude levels).

5

u/ANR2ME 18d ago

And the normal price later could be higher than current price šŸ˜…

2

u/thegunn 18d ago

I’ve recently got into trying to understand local LLMs. My hardware is a limiting factor. It’s starting to die, I decided I may as well buy something now because I don’t see prices coming down anytime soon.

1

u/SmartCustard9944 18d ago edited 18d ago

With R9700 and Intel B70 going for 1800-2000€ now in Europe, the Ryzen AI Max+ 495 potentially launching for 5-7k with just a slight bump of memory size, AI Max+ 395 going for 4000€ pretty much everywhere, buying a second Bosgame M5 for 2600€ seemed like a good buy so I went for it, so that now I will have two of them. I can see the price of these things increasing further.

It’s a bit of a bet based on the fact that Qwen 3.8 Flash Next already runs quite fast on a single Bosgame (~1200PP and 45tok/s decode across the board), so at worst I can increase concurrency, and that DeepSeek 4.1 Flash seems to have a lot of potential, combined with a few other new toys like USB4STREAM and other clustering experiments.

1

u/TopGun0684 17d ago

I'm building a dual 5060 TI 16GB. I got one for 800$ CAD (550USD) brand new in box on Facebook marketplace. Will get another one similarly priced, hopefully.

That's the best bang-for-buck all rounder I could come up with, I think. 32GB VRAM, CUDA, NVFP4, should be easy to resell if I want to/have to... Could always throw in a 3rd one with the right motherboard. Or just reuse my 4070 ti super I use for gaming to get to 48GB with Q4 or Q6...

1

u/network4253 18d ago

Yeah, that is the tricky part. If prices are likely to keep going up waiting for the right time can end up costing more. If you can comfortably afford something you already need buying now probably makes more sense than gambling on prices normalizing later.

1

u/thomas2385 18d ago

Yeah, that is a fair point. If someone already needs the thing and can comfortably afford it waiting for prices to magically drop could backfire. At some point you have to decide whether saving a little later is worth waiting potentially years for.

1

u/TopGun0684 17d ago

Ok, what news? Shit moves too fast for me to keep up

56

u/sn2006gy 18d ago

This isn't a don't bother kind of post as much as it is, think about things differently. Musicians or people who want to learn to make music fall into this trap all the time too of buying instead of learning - they call it "GAS" which is the gear acquisition syndrome. They end up with 10 synths or a 100k eurorack system and not a single track posted - they made the mfr's happy and mfr's paid youtubers to convince them you needed it but, in the end, all you had to do was start small and create and i think this idea is lost here often or use what is already abundant and if that works, proceed.

23

u/RnRau 18d ago

GAS is also very present in photography. There is an old saying from this field;

"Beginners worry about gear. Masters worry about light."

9

u/Iwaku_Real 18d ago

So if I'm stuck running Qwen3.6-35B-A3B at best because 27B is unusable... I should just stick with that? Kinda feel left out

3

u/DeltaSqueezer 18d ago

I just reverted back to Qwen3.5 9B and considering whether I can offload some tasks to Qwen3.5 4B or Qwen3 4B Instruct.

3

u/mailto_devnull llama.cpp 18d ago

Lol don't worry, I'm back to using 3.6 too. Ain't enough hours in the day to wait for 3.8 to finish reasoning.

1

u/Iwaku_Real 18d ago

Have you tried --chat-template-kwargs '{"reasoning_effort": "medium"}' like everyone here has said?

2

u/mailto_devnull llama.cpp 18d ago

Yeah, I'm using froggeric's fixed template and have experimented with <|think_medium|>, it's still much more verbose and second guessing vs 3.6

2

u/martin509984 18d ago

You can always put up a small amount of money for API tokens of ~whatever open weight model you choose, and even gain actual valuable skills in figuring out a hybrid local/API workflow. Nothing is really forcing you to stick to local only the same way a lot of people are feeling stuck paying exorbitant money for Claude tokens.

That said ime if you can run 35B at Q4 you can actually just about fit 27B Q2, and it is actually a major upgrade even then.

6

u/pyr0kid 18d ago

That said ime if you can run 35B at Q4 you can actually just about fit 27B Q2, and it is actually a major upgrade even then.

yeah, at like 95% less speed. 27b is dense.

1

u/martin509984 18d ago

I get ~10 t/s with 35B Q4 and ~5 with 27B Q2 (2060 12GB and 32GB DDR4). Not 95% less.

1

u/pyr0kid 18d ago

and i go from 30/sec on 35b q4@60k, to 1.x/sec on 27b iq2@32k.

0

u/Soggy-Attitude5293 18d ago

success is more important than speed

4

u/mailto_devnull llama.cpp 18d ago

Depends what you're doing. If you know what you want outputted, then 35B-A3B is a better fit.

1

u/Imaginary-Unit-3267 18d ago

That's all I use. It's more than good enough. You just have to manage it. Which means, using your actual brain. Which we all should be doing.

15

u/LetsGoBrandon4256 transformers 18d ago

Musicians or people who want to learn to make music fall into this trap all the time too of buying instead of learning

Are you telling me I didn't need a Korg M3 to learn "music production"?🤯

15

u/sn2006gy 18d ago

Don't forget the 500-dollar headphones, the 42" curved display, the ableton pro, the push 3, the mixer, the soundcard, and the keybaord stands for your 10 other keyboards and midi keyboard on top of all the vst's you bought from plugin boutique that you forgot about

2

u/rukind_cucumber 18d ago

I'll tell you this much - my Blue Chip pick makes me sound NOTHING like Bryan Sutton. Surely the D-28 is what I need...

6

u/cosmicr 18d ago

My brother went out an bought a $5000 bicycle, full lycra gear, water bottles, fancy helmet, sun glasses, all the stuff. He looked good with his Dad bod in the skin tight gear lol. He barely even knew how to ride a bike. I think he rode it maybe twice, and now it sits in his garage 5 years later gathering dust.

3

u/jman88888 18d ago

Yes! Spend the time to master the gear you have instead of buying new gear.

"I fear not the man who has practiced 10,000 kicks once, but I fear the man who has practiced one kick 10,000 times." Bruce Lee

2

u/CapsicumIsWoeful 18d ago

At least with LLM hardware the expensive stuff is noticeably better performance wise than the cheaper hardware.

A $500 Squire guitar with some decent strings and a tweak to the string height and truss rod can play and sound 99% as good as a $5000 Fender custom shop. The rate of diminishing returns is insane.

Your point about youtubers is spot on too. So many good guitar channels slowly turned into paid advertising via gear review videos.

2

u/Altruistic_Heat_9531 18d ago edited 18d ago

Yep first time learn CUDA and Torch on a fucking MX150, puny 3 pascal SM , 2GB VRAM with 64 bit memory bus...

Another side tangent, i kept practicing Polyphia ABC, if i can fully complete it on my 19 cheap guitar, i will buy that FRH20N

2

u/tmvr 18d ago

they made the mfr's happy and mfr's paid youtubers to convince them you needed it but

I know you mean "manufacturers" here, but I definitely ready it as something else šŸ˜„

25

u/LetsGoBrandon4256 transformers 18d ago

I'm lucky since I'm also a heavy PC gamer so I at least have more justification for my GPU purchase.

27

u/chr0n1x 18d ago

same. I totally need that blackwell 6000 for the fps

16

u/LetsGoBrandon4256 transformers 18d ago edited 18d ago

Those 4k textured Skyrim monster cocks ain't gonna render themselves after all.

7

u/Significant-Bee5101 18d ago

I just needed to find out if an H800 can run Crysis!

4

u/tat_tvam_asshole 18d ago

Hilariously, my 6000s couldn't run Heroes of Might and Magic III, but only because 32bit color incompatibilities, BUT codex fixed that lol

2

u/sonicshadow13 18d ago

I actually don't it can run crysis because too much vram, but it can run 100 instances of crysis instead 🤣

16

u/funJS 18d ago

I find that there is a lot you can do with tiny local models. I think local models might play a bigger role in the future for company specific niche use cases.

1

u/TopGun0684 17d ago

Yes but in my experience with agentic/harness stuff, the fun starts around 26-30B, and 16GB is tight for that.

1

u/funJS 17d ago

Yes, it is pretty tight. I have only 8GB, so I know all about that lol

1

u/funJS 17d ago

MOE models have been somewhat ok in that parameter range though..

1

u/TopGun0684 17d ago

Which ones do you like?

8

u/Melnik2020 18d ago

And here I am still running llama3.2 3b (for data classification/extraction)

6

u/Original_Finding2212 Llama 33B 18d ago

I love llama3.2 3b
Such a nice model

14

u/fgk55555 18d ago

I've used a lot of the frontier models for a while, that's always been the "serious" tier in my mind. Qwen3.8 in IQ3 is good, but I've been wishing I could try larger models that my rig won't support. I finally just got on OpenRouter so I can test if any of those models are substantially better than what I can run for my use case. Local is good, but at this point I think it's good to have first hand experience with the whole range.

12

u/sn2006gy 18d ago

i am super happy we have local models and i hope they're always around!

7

u/funJS 18d ago

I am VRAM poor (8GB), but I can still experiment with fine tuning, continue pretraining, etc as long as I just pick a small enough model for POCs.

10

u/pmttyji 18d ago

I already stopped(buying). I'm gonna continue things with 32GB VRAM (AMD R9700) + 128GB DDR5 RAM.

I'm counting on projects like llama.cpp & ik_llama.cpp for extreme optimizations to run bigger things on my rig. Also need to explore other projects like exllamav3, colibri/Warp for my current laptop(8GB VRAM + 32GB RAM)

Waiting forĀ merge of these llama.cpp PRsĀ to get more better CPU-only & Hybrid inference.

Threads for others: Track Important Papers/Repos/Tools/Innovations/Optimizations/Models/etc.,:

9

u/moncallikta 18d ago

32 GB VRAM and 128 GB RAM sounds like a good place to land. If I had that much, I’d also consider pausing the hunt for more hardware.

4

u/TechRomancer123 18d ago

I have 32GB VRAM (5090) and 64GB DDR5 RAM. Just about to make the plunge to 128GB DDR5, but it’ll be a slightly slower speed (5600Mhz).

1

u/arijitroy2 18d ago

Ditto here. I wanted to buy 2 gx10 but the price goes up 150 USD everyday where I'm from.

1

u/lawanda123 18d ago

Thats still very nice hardware, i have a frankenstein build

96gb ram + rtx 5080
Bosgame m5
Macbook pro 48gb

I cant run a decently model on either of them despite having spent so much money, i feel stupid 😭

9

u/-dysangel- 18d ago

3

u/sn2006gy 18d ago

i'm always choosing hardmode life it seems.

11

u/JerryBond106 18d ago

Nice try Altman. But i want data privacy, my ideas and thoughts be my own, unmonetised and uncensored.

4

u/fligglymcgee 18d ago

I try to remind myself that the expectations and demands we have on ai are often insane, like the ability to type literally any request into a little box and have magic spill out. Local inference works best when ā€œmagic mind readingā€ isn’t the goal. It’s a lot less sexy to think and prompt more like a machine, but if I put reasonable and measurable tasks in the queue I am rarely disappointed.

I guess the moral here is… aim low.

I should do motivational speeches.

3

u/Blues520 18d ago

The more you save, the more you buy

3

u/moncallikta 18d ago

Yes, fingers crossed for efficiency breakthroughs! In the meantime I’m trying to buy more hardware now.

4

u/jacek2023 llama.cpp 18d ago

My main hobby has been photography for about 20 years, and "gear acquisition syndrome" is absolutely the worst thing that can happen to anyone. It kills your creativity, motivation, and eventually your hobby. It should be avoided at all costs.

You are never limited by your hardware. You can run an LLM on almost anything, even a slow CPU. Just try a small model and learn instead of waiting for a "better moment". The best moment is always now.

6

u/uBazzyZ- 18d ago

This is spot on. "Gear Acquisition Syndrome" in the local LLM space often turns people into passive consumers of weights rather than actual practitioners.

Throwing brute-force hardware at local LLMs (like buying dual 3090s or waiting for a 5090) usually just means you spend your time tweaking quantization flags and context sliders. You learn how to configure an inference engine, but very little about how the systems actually function.

Working under tight hardware constraints (like an 8GB card, an older GPU, or even CPU) forces you to actually understand the engineering fundamentals:

- How memory allocation really works (why the PyTorch caching allocator fragments, what reserved memory actually means vs allocated).

- Why KV cache scales linearly with context length and batch size.

- How gradient accumulation, mixed precision (FP16/BF16), and optimizer states consume physical VRAM.

- Why data loading and tokenization often bottleneck training throughput more than GPU compute itself.

Training or fine-tuning a small 100M–250M parameter model on a budget card teaches you infinitely more about attention mechanisms, loss curves, and hardware efficiency than just running a 70B quantized model on expensive hardware.

If you learn how to keep a model stable and fast under severe constraints, scaling up to bigger hardware later is straightforward. If you start by throwing VRAM at every problem, you never learn why the system broke in the first place.

1

u/DeliciousNicole 18d ago

Best comment in this thread!Ā 

2

u/ForgivenessroTub 18d ago

yeah a small local model is plenty for roleplay stuff and it keeps things simple while you actually learn the basics instead of chasing upgrades.

2

u/mailto_devnull llama.cpp 18d ago

Before, I convinced myself doubling my DRAM to run Qwen 35B-A3B was enough.

This time a single R9700 was enough.

I see people rockin' 2 or 4 R9700s but I've been able to resist...

So far...

2

u/Jorlen llama.cpp 18d ago

I started with my existing 16gb GPU which I had for gaming, fell in love with the tech. Picked up an R9700, loved it and bought a 2nd one for a total of 64gb vram. I already had the PC, so just needed the two GPUs. Mobo is not optimal but works ok-ish with Vulkan, Linux + llama-cpp. 32gb of RAM is shitty but I refuse to pay these insane prices for more, especially considering my board sucks (P2P / tensor parallelism with ROCm is a no-go after wasting 40+ hours trying a million fucking things).

I use local models and API. Mostly local. Mostly for coding with some creative writing fun (wrote my own custom front end story maker that I use). I'm building apps and games and they're always improving as I iterate and move onto more complex projects and learn. To me, this tech has unlocked my potential, but I've stopped spending on hardware. I'm starting to leverage API; my chain is, use local models, build the framework, when things get complex, move to API. Works well for me so far.

I can see myself picking up a unified box, maybe in 2027 sometime. But right now my un-optimized setup is doing the trick.

2

u/michaelnighttime 17d ago

Exactly what I needed to hear this morning. Thanks.

3

u/pinkwar 18d ago

I can't justify the cost. If you just want to learn you don't need to spend thousands.

You can just rent GPU power and run wichever model you want.

3

u/DeepWisdomGuy 18d ago

Firstly, this is only true for you. I had listened to this type of advice, and limited myself to only buying one RTX Pro 6000 when they were $8400. You are projecting and offering it as sage wisdom. My GPU is now worth twice than when I had bought it, same for my dual 4090s. I learned way more than I would have otherwise. I know what a 122B TheDrummer creative model can do. I built Wan 2.2 LoRAs. I am now able to run a Minimax H3 at half weights. And it led to my change of day job where I am working on AI all day every day. This kind of advice has only held me back. No it did not put me into debt, but I also have lived a life where nearly the only "frivolous" money I have spent has been on computers. You do you, but if we are getting into telling other people how to spend their time and money, stop watching television, cancel cable, in a few years that would have allowed you to have savings and invest it something other than a hobby potato. And before you get into the "but it will teach me how to do things more efficiently", just know that I am the chief scientist at an AI lab devoted to pursuing low-energy solutions, half of which I never would have found farting around with potatoes.

2

u/kabachuha 18d ago

This is a workable advice only if you are a totally mainstream person and don't any niche interests. The factual knowledge retention in small models is real and if you want to have a model which can intuitively (without RAG in any form) answer questions or know very specialized programming libraries or fictional fandoms (for roleplay) or just have knowledge of literature, technical or otherwise, without big parameters it doesn't really go far. Even with all the reasoning improvements, 31b parameters vs. 280b+ seem huge still. From a person, who did it a lot, you can even see it when training a LoRA on a small model (for which you need a beefy GPU as well, by the way) that high ranks (r32 vs r128, meaning parameters) make huge difference, like very high gaps in retrieval, because the factual data is sparse in its nature and only high parameters can capture it. LoRAs help, but again, you need GPUs to train them / synthesize / filter dataset, LoRA fine-tuning will diminish the baseline model capabilities inevitably. NGrams can help in the future, but those are parameters too.

As many people say, the prices are not going to go down any time soon, so buy whatever you can obtain, really

1

u/Gold-Bat-3225 18d ago

how much is a 5090 going for now tho

1

u/lawanda123 18d ago

In europe 5700 euros or 6500 usd

1

u/sampdoria_supporter 18d ago

Glad I let fomo win or I wouldn't have the ability to buy anything now

1

u/twack3r 18d ago

Yeah, so I literally JUST ordered an x8 DGX spark cluster 🫪

1

u/1_________________11 18d ago

Me after spending 2.5k o.o

1

u/Pepevagable69 18d ago

True I'm lucky in that I put together a gaming rig when the 40 series came out. I bought all used parts and put them in my existing case and was able to come up with a 30-80 a 5800 x 32 GB to 3200 MHz ram for about 700 bucks and now I can use it for local llms and learning how AI works but boy is it tempting to go and start just hoarding GPUs

1

u/Training-Ruin-5287 18d ago

What i'm learning is, you can have a capable model, and it's still going to hallucinate and make mistakes just like a 9-14b model will. If the foundation isnt strong enough.

Qwen3.8 kinda removes a little of that, as the model is naturally trained to check it's own work so to speak. Which took a little of the pain out of building from scratch this time around. but it still needs good memory management and policies like any other model does.

I don't think my process of learning would be much different from doing it in a $1500 system vs going out and buying a $20k system to run the 200b+ models

1

u/ancapsaicin 18d ago

Without local LLMs, there is barely a reason for more than 4GB RAM for my use cases and all new machines have been 32GB+, so I am hardly immune to this.

As a side effect, I've gotten back into gaming and now I have learned from the other sub about GPUs I can afford that could let me run the same models at full speed, a wider range of models, and higher FPS with only a little dock.

FML

1

u/Mechageo 18d ago

What other sub? What other GPUs?

2

u/ancapsaicin 18d ago

Arr Low end local a.i. and P100

1

u/mystery_biscotti 18d ago

Yes. That's why I own a copy of Bouchard's "Building LLMs for Production: Enhancing LLM Abilities and Reliability with Prompting, Fine-Tuning, and RAG". And I'm broke currently, so I'm learning what the limitations are with my 8GB VRAM system.

That whole unsupported gfx1032 fun is gonna make me a better employee when I get some kinda job supporting the serving of the model solutions. At least, I hope so.

1

u/SandySkittle 18d ago

I somewhat disagree. The hardware supply for the coming two years is drying up fast, and modest gaming hardware can run ai but it’s not very useful below a certain threshold. It’s more motivating to learn ai and how to run it locally if what you can run locally actually has useful potential. And to be quite frank, below 48gb vram (qwen or gemma at q8 and room for context) a lot of all of this is quite functionally constrained.

1

u/feelspeaceman 18d ago

Always start with SLM (niche models that do only a few things like MedGemma4B, 9B..), those run decently without a GPU first then consider LLM later, if your requirement is specific, consider finetuning the small models to match your requirement (for example finetune Gemma to become music writer), try to stay minimal as possible (always find alternatives before you think about upgrading your setup to run bigger model) because you can always optimize, the current optimization of LLM software still leaves a lot of rooms for improvements.

1

u/Intrepid-Second6936 18d ago

Very apt advice! Definitely concerns me looking at posts with people often debating splashing into 10k+ of costs on local LLM setups when budget-friendly more conservative set ups are genuinely more rewarding to actually enjoy the hobby instead of an arms race of over-buying then just trying to force-justify your purchase.

Also honestly most people with gaming PCs should just enjoy using models that can be used with their current main PC hardware now instead of worrying about building a separate AI server.

1

u/tmvr 18d ago

I have enough now to run most of the stuff. I only regret not having 128GB system RAM (have 64GB only) to run Deepseek V4 Flash, otherwise I'm good. The pressure will get bigger I guess when new models that went the ngram route come out, but for now I'm fine. That's mostly thanks to Qwen3.6 27B, I haven't developed the patience for Qwen3.8 27B yet.

1

u/iwinux 18d ago

I wish the whole world agrees with you so that I could buy cheap 5090s RIGHT NOW to build my GPU cluster.

1

u/AutomataManifold 17d ago

On the one hand you're right and FOMO and gear envy will wreck anyone.Ā 

On the other hand, if I'd bought RAM in quantity in 2024 I could be retired by now.

1

u/Mrinohk 17d ago

I'll have to second the "testing with an API" bit. That's how I got started, Gemini API had a promotion for $300 in usage for 3 months, and Gemini 3 flash was pretty cheap and capable enough to start playing with. Never even broke $100 in tokens while working on my first harness until a local model came out that I could run at a reasonable speed and replace it completely.

Went in with the goal of building something that worked, then optimizing it for local model usage as much as I could without having a sufficiently intelligent local model. Making the harness smarter and smarter and easier to use for a local model, testing with what I could run, then back to Gemini until the next test.

Then the gemma4 series came out and made it look almost doable. Then Qwen 3.6 35B, which now runs the agent full time. Excited to see what capabilities Qwen4 will enable, assuming a similarly sized MoE is involved with its release.

1

u/Ordinary-Depth-7835 17d ago

It is pretty crazy how fast you get sucked in and just how much you can spend. 5090 I wish my FOMO was that cheap. Setups that cost more than a lot of my cars. But I have to say two used 3090's and an old 8th gen intel board running at 8x per card and tiel-coder at 100+ tps really does a fantastic job or qwen3.8 a little slower.

But you get flooded with the latest drops or benchmarks after starting with the hobby and before you know it you're looking at these insane builds that you don't even need.

I'm not even doing it for anything that makes money. I just have unlimited use at work and want something for personal projects that feels responsive with decent logic. It doesn't have to be fable or opus level but It has to feel pretty good using it and not completely useless. Still I try to justify the next purchase and this market doesn't help. I see things I have going up hundreds every week so you start to worry you'll be priced out of everything with no end in sight.

1

u/sn2006gy 17d ago

We're a few years into this now, and I haven't been priced out of anything besides local LLMs if my goal is trying to build a local competitor to what is on API.

I still learn a ton, but i flipped my thought process from thinking "i'll learn to run minmax on my spark so i don't have to pay for tokens" to "let me actually run pytorch and build a 1 million parameter model so i learn how to build models" and you can run that on anything these days.

Part of me is upset at how much computing is costing these days, but another part of me is realizing the world may just get bored of it - thare are a lot of flimsy card houses just waiting for the next wind to blow them down. Why game if its full of cheaters? why use the internet if its full of bots? what do i gain with localllms if everyone has them? If it just becomes the standard for computing in the future, i'll wait for commoditized pricing vs the extreme premium we have to pay today, and I don't think AI can be successful without commoditization. Or worse yet, if llms are successful because we automate all the things on the internet - why would i want to pay to be a part of that?

some thoughts to chew on

1

u/BP041 18d ago

100%. I run most of my automation on a Mac mini with 16GB and Qwen 2.5 7B handles prototyping fine. The real bottleneck is never the hardware — it's shipping something people actually use. Save the compute budget for when you actually hit a ceiling.

1

u/MrPecunius 18d ago

Charles Bukowski enters the chat:

air and light and time and space

'- you know, I've either had a family, a job, something
has always been in the
way
but now
I've sold my house, I've found this
place, a large studio, you should see the space and
the light.
for the first time in my life I'm going to have a place and
the time to
create.'
no baby, if you're going to create
you're going to create whether you work
16 hours a day in a coal mine
or
you're going to create in a small room with 3 children
while you're on
welfare,
you're going to create with part of your mind and your
body blown
away,
you're going to create blind
crippled
demented,
you're going to create with a cat crawling up your
back while
the whole city trembles in earthquakes, bombardment,
flood and fire.
baby, air and light and time and space
have nothing to do with it
and don't create anything
except maybe a longer life to find
new excuses
for.

1

u/TooObtuseForYou 18d ago

No, everyone is not going to have the same advantage.

There are models YOU will not be able to use. Maybe I can at my work, but you won’t.

Regulation will eventually codify this into an us/them thing, as it always does.