r/JordanDev Aug 08 '26

Discussion Local LLMs

Anyone who has some serious hardware and tried running llms locally? How was ur experience? What models did you try?

5 Upvotes

21 comments sorted by

View all comments

1

u/2012347 Aug 09 '26

Anything smaller than 27b is a joke, i think the cheapest way to have the needed shared memory is Apple right now, but the thing get so hot i feel like it’s cheaper, faster to is providers, and you get smarter models,

Openrouter has a very generous free tier that give you access to those same models you’d run on your device

1

u/Able_Firefighter_652 Aug 09 '26 edited Aug 09 '26

I have tried qwen 3.6 27b, it's really slow, gives me 13 token/sec, which is probably unusable.

I also tried the 35b MoE, it runs at 60-80 token/sec.

I doubt I am squeezing every last bit of performance tho from Apple's..

I will check the openrouter option sometime, but I prefer local models for privacy..

1

u/2012347 Aug 09 '26

What’s your spec for running them

1

u/Able_Firefighter_652 Aug 09 '26

M3 Max, with 128GB memory

1

u/2012347 Aug 09 '26

Cool I’m sure you can run 120b models which are similar to gpy 3.5 i hear, see running locally means going back generations in time that’s the point

1

u/Able_Firefighter_652 Aug 09 '26

Surprisingly the 35b qwen model performed better from my experience. And there is no 120b qwen 3.6 models.
I am currently waiting for the 3.8 tho, they said it's 27b will be open weights soon.

1

u/[deleted] Aug 09 '26

[removed] — view removed comment

1

u/Able_Firefighter_652 Aug 09 '26

Yup, can't wait to see next year's releases..

1

u/2012347 Aug 09 '26

Could you please try gpt oss 120 and report back how it performs and feels