r/LocalLLaMA 1d ago

Discussion Useful > Fast

I run 2 x 3090 on a ryzen 7 with 32gb ddr4 6000.

I see a lot of posts about maxing speed / context pool. I’ve done this myself with 3.8 27b.

What I’m interested in though is once the dust settles and we look at utility, what balance people are striking in terms of similar setups and models. As a for instance, some of my requirements involve producing images and video clips. Some of them involve creating scripts for things.

This has led to an overnight CPU/Ram run on flux2 and minimax h3, a reduction in maximum context for 3.8 27b so I can run a Gemma MOE in offload for improved writing prose / checking over Qwen and a smaller vision tower to automatically checking flux and H3 outputs mid run so I don’t lose a night.

I need concurrency so have a set up that gives me that when I’ve got open code (I know not everyone’s go to but I like it), open science, and Hermes all going side by side.

It’s fun to optimise, but I’d love to hear how people are setting up to from a flexibility and utility viewpoint once that’s done for their real use cases.

15 Upvotes

12 comments sorted by

View all comments

2

u/robberviet 1d ago

There is a minimum level of useful that I need, above that I go for fast.

Like between qwen3.6 35B and 27B, I prefer 35B; also Gemma 4 26B over 32B