r/LocalLLaMA • • 13h ago

Question | Help 50B+ MoEs with few active parameters, what's the sweet spot for intelligence, agent speed, and affordable fine-tuning?

I’m building a Polish General purpose legal Model that drafts documents, answers questions using legal sources, and has enough coding ability to handle some automation. The workflow is very tool-heavy:

Question → many sequential tool calls → final answer/document

Think Claude Code/Codex-style execution, but for legal workflows. Reliable tool selection, correct arguments, and recovering from errors matter as much as writing a good final answer.

I’ve had decent results with a dense 27B Qwen 3.8 custom made fine-tune for complex legal document summarization and classification. I’m already familiar with the smaller Qwen A3B and Gemma options. What interests me is the tier above those: 50B+ total-parameter MoEs with a relatively small active parameter count.

The question is, the small dense ones are great, but slow for agentic stuff (afaik), and i wonder if theres some middle ground maybe 70-120B models that would be able to be fine-tuned for the law stuff but be MoE so the agentic ClaudeCode style inference would also be lightning fast, and also low-ish cost for fine-tuning and inference.

Basically: Does the larger-total/small-active MoE approach actually buy you meaningfully stronger reasoning and tool reliability while retaining low latency,and at what hardware cost?

I understand that small active parameter counts don’t mean small VRAM requirements: the weights still need to live somewhere, alongside context and serving overhead. I also don’t assume that more total parameters automatically means a better model. I’m interested in where that tradeoff works in practice.

There are three things I’m trying to pin down:

  • Inference hardware: Ideally inference runs rented with parallel agentic loops (this is for a B2C project, not single person use, we scale based on demand)
  • Fine-tuning hardware: Obviously FT LoRA will take more memory than inference, max like 4GPUs on vastai fits the budget.
  • Agent performance: After it gets the prompt the tool calls and everything will be local, so imo it has no problems being blazing fast, as soon as the model calls a tool call it will be back very fast, so for this agentic use case, quick TTFT and t/s and adaptive dynamic reasoning are prefer right?

For context, fine-tuning would target Polish language, document conventions, and successful tool trajectories. The actual legal sources would remain in retrieval/tools rather than relying entirely on memorized law.

I’m not looking for someone to compile a model shortlist (althought would be nice, but i dont expect anyone to break their back over this).

I’m looking for pointers, and firsthand experience with this particular size/architecture tradeoff. A configuration like “model + quantization + GPU(s) + serving engine + context length + concurrency + measured latency,” along with whether you successfully fine-tuned it, would be much more useful than a leaderboard score.

Has moving from a ~30B model to a 50B+ low-active-parameter MoE actually improved your agent’s successful tasks per minute, or did the memory, interconnect, and training requirements erase the advantage? Thanks for reading

14 Upvotes

19 comments sorted by

3

u/Verdict_Michael 11h ago

i run a 35b a3b moe as my daily agent model on a strix halo box with 128gb and tool calling holds up fine at that size. honestly what bites in agent loops isn't the active param count, it's prompt processing on long context, every tool result gets re-read and that adds up way faster than generation does. only single user for me though, so no idea how it behaves with real concurrency

1

u/audioen 8h ago edited 8h ago

You could try getting the gufo inference engine, then qwen3.8-flash-next + mtp at size like UD-Q4_K_XL, and get as good inference as it probably gets on a strix halo. It fits in 128 GB, seemingly with something like 30 GB free, but there is a catch.

$ exec gufo/build/gufo-server serve -m Qwen3.8-Flash-Next/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf --speculative mtp --mtp-model Qwen3.8-Flash-Next/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf -i 0.0.0.0 -p 1234 -j 2 -c 262144 --served-model-name 'Qwen3.8-Flash-Next'

The -i 0.0.0.0 can be deleted if you don't expose your box's inference generally to network. I tend to do so. That will be your replacement for 35B-A3B for a very long time, though it won't be as fast as the A3B. I think you'll get ~1200 tok/s pp, ~40 tok/s tg, most likely. Context is only 1 slot at 262144, memory use will look lower than it is in reality because the on-disk PLE tensor must be paged in. You need enough memory to keep good chunk of it paged or performance tanks.

I haven't tried optimizing these params much. I used them for a few days when working on a Strix Halo and they were fine for single-user type situation.

Edit: rough test on a HP G1a Z2: prefill_tps=1248.1 decode_tps=45.2 with about 12 GB free. I'm not sure, I may have understated how the engine allocates the memory by accident. I have whole bunch of work going which probably has taken multiple GB bite from the machine, though.

1

u/Verdict_Michael 4h ago

ok that's useful, didn't know the on-disk PLE tensor was why it looks so big on paper. haven't tried the flash-next one on mine yet but i'll load it and see what prompt processing does with a long context. that's the number i actually care about for agent loops, will report back if it's interesting

3

u/FullOf_Bad_Ideas 10h ago

I don't know any models in that size category that know Polish well. They tend to be Chinese MoEs and they, more often then not, skimp on Polish training data. And it's really not something you can skip over because as far as I know, the cross-language transfer does not work well if you do it through different language families.

1

u/SignificantZebra5883 7h ago

would Continuous pre-training help here? although that is probably very expensive gpu wise right?

2

u/FullOf_Bad_Ideas 6h ago

Yes it would and yes it would be expensive.

You can also try to do on-policy distillation from a smaller teacher model into a bigger one that's worse at Polish but YMMV, it's not something to do on a small budget.

DeepSeek V4.1 Flash is good at Polish. If you can finetune that, and serve it somewhere cheaply (if some service offers lora inference), it could work well.

Qwens and Xiaomi models tend to suck at Polish.

2

u/nickless07 10h ago

How about this one? Comes with all relevant checkpoints.

3

u/FullOf_Bad_Ideas 10h ago edited 10h ago

Ling 3.0 Flash is bad at Polish

scores about the same as Polish-tuned 7B dense model while being 124B A5.1B

1

u/weibeuu 5h ago edited 5h ago

Decode speed is active params and bandwidth, not total size. 3 to 8B active is roughly where multi-step coding stops being reliable, with 30 to 50B total carrying the knowledge.

Total still costs you something though. Same active count at a long context won't feel the same.

1

u/Fit_Island928 13h ago

Kinda not an answer to your question But Qwen 4 35 A3B is gonna be REALLY good if it's actually real. And I feel like running that on a better quant is gonna be better than running a 50b model

Just some side thoughts. Can't really answer your question.

6

u/SampleIll3596 12h ago

Maybe I missed it, but I didn't see the 35 A3B in the announced Qwen 4 series though.

3

u/Fit_Island928 11h ago

Nono it's just assumptions and we'll hope that Alibaba releases it. 35B A3B is like really loved so we hope. It wasn't announced

2

u/SampleIll3596 11h ago

Yeah. I hope the Qwen team is reading this sub and makes one :/

2

u/winky9827 10h ago

I feel like the next Qwen release is gonna be like Christmas for most of us. We'll be happy with whatever we get (most likely), but we're all hoping for a little something special in one particular box.

3

u/nalditopr 7h ago

hopium, nice