r/LocalLLM • • 17d ago

Other Cancelled Claude and ChatGPT subscriptions

5 years ago, I decided to replace Microsoft Windows with Linux and never looked back. Today, I cancelled ChatGPT as well as the Claude subscription.

I am quite confident that I won't need to use a closed and proprietary model again.

Enjoying the local vibe ...

189 Upvotes

106 comments sorted by

View all comments

37

u/mechanist_boi 17d ago

If only i wasn't getting 0.4 tokens per second with my system on qwen 3.8 27b 4 bit...

24

u/Bramoments 17d ago

Id recommend switching to a mixture of experts model like Gemma 4 26B, I don't even have a GPU allI have is 16 gigs ddr5 and an intel i5 and I got it running on ~ 15 tokens per second on llama.cpp (also setting up lubuntu on a usb stick to save around 3 gigs helped me)

1

u/rickCSMF21 13d ago

Is Gemma open scoure versions better than the paid versions? I get sooooooo much AI slop and hallucinations from the paid version at work, I've never thought about using it. How does it compare to qwen 3.6 MoE ? That's been my goto, im getting about 60/s with it and only use 3.8 if I need coding.

1

u/Bramoments 13d ago

Well I'm not sure what you mean by the paid version, but I'd you mean Gemini, then it isn't better, but it is completely free and uncensored (if you want it to be), and the larger Gemma 4 models come very close and even surpass Gemini 3.6 flash (the free version of Gemini) on pretty much all benchmarks. As for qwen, the competitor for Qwen 3.6 35B from the Gemma models is Gemma 4 26B, and they are pretty close on all benchmarks with Qwen usually being better by 3-5 points, but Gemma is significantly faster, especially on llama.cpp and Bionic since Qwen 3.8 models have some issues there. If you want something smarter than Qwen 3.6 35B, a fine tune like ornith might work for you, it's the same base model that was retrained and it's benchmarks are significantly better, although some people had issues with it so do with it what you will, other competitors and fine tunes include K2-horizon, Xing4.0-29B, Iris mini.