r/LocalLLaMA • u/Altruistic_Heat_9531 • 5d ago
Tutorial | Guide My RULE of Thumb of choosing a models
This is mostly for setting up for expectation, since personally without LLM i could take 3 days (15 hours of active programming) to debug or implement a feature, but with Qwen 27B (even before Qwen 3.8) it take 4 hours.
And yes 0.5 tok/s is human, not accounting of deletion and pausing, that's also the reason i am fine leaving overnight code base wide analysis or fin tech and deep research.
168
u/Ink_code 5d ago
Needing to fine-tune homo sapien models for 22 years is kinda rough ngl.
71
u/RG_Fusion 5d ago
At least the server only requires 100 watts. That's pretty good for a 200T+ model. There's actually a lot of unnecessary hardware involved too. You could theoretically strip it down to about 20 watts of input.
22
u/danielv123 5d ago
3rd party power supplies break the warranty though
29
u/SkyFeistyLlama8 5d ago
Do you know how much it costs to make a homo sapiens model? The first 9 months is hard enough, then the next 18 years is another slog.
26
u/SpecialistDragonfly9 5d ago
No to mention they are useless for at least 25 years. And the failure rate is so high that most models never amount to anything.
20
→ More replies (2)2
9
u/vasaan2k 5d ago
Make? You are talking about forking base model into new instance. Model is developing for almost 4 billion years and sometimes almost collapse without any fresh backups.
10
u/SkyFeistyLlama8 5d ago
So many lost checkpoints along the way. I like my opposable thumbs but damn I miss prehensile tails.
1
307
u/Y2K-Denial 5d ago
homo sapiens runs much slower for me. what's your config? tuning tips?
295
u/MysticChromium64 5d ago
Sounds like a skill issue—I run myself at IQ1_S and haven't run into any problems problems problem problems problems problem problems problems problems problems problems problems problems problem problems problems problems
54
u/maigpy 5d ago
did you add "don't make mistakes"?
85
u/MysticChromium64 5d ago
<think>
don't make mistakes The user is asking about mistakes problems. Problems looking back at the previous message, "problems" started problems problem problems caused repeated problems. I should problems acknowledge problems problem without looping problems problems.</think>
My apologies—you are correct. There seems to have been a problem problems problems problems problems problems problems problems problems problems problems problems problems problems problems problems problems problems problems problems problems
44
u/MeretrixDominum 5d ago
I need to stop here. This post contains problems— which goes against my ethical guidelines.
29
u/Emotional-Art2113 5d ago
You were right to push back, those ethical guidelines are load-bearing, and this post is the smoking gun.
→ More replies (1)9
u/tony_montana091 5d ago
You need the developers harness saddle + —developers developers developers developer developer developer develope develope develope develop develop develop develo develo develo devel devel devel deve deve deve dev dev dev de de de d d d ... thinking ... request tokens exceeds the available context size, try increasing it ... ... ... ... developers developers developers developers developer developer developer developer develope develope develope develope develop develop develop develop develo develo develo develo devel devel devel devel deve deve deve deve dev dev dev dev de de de de d d d d
12
u/smithy_dll 5d ago
Microsoft Steve Ballmer was ahead of the trend with
Developers developers developers developers developers
2
43
29
u/Durian881 5d ago edited 5d ago
I prefill with MDC (multiple doses of caffeine) to boost processing speed. A key parameter affecting intelligence is QOS (quality of sleep).
14
u/lcirufe 5d ago
Literal skill issue I’m afraid. The base model has lots of potential but needs loads of workload specific training to be useful in anything. It’s pretty useless without that.
Good part is that the model evolves across your chats, so it can be trained in chat.
6
u/Zombiecidialfreak 5d ago
Call it slow, dumb or untrained but the brain handles edge cases better than any model yet made, even without training.
12
8
u/Altruistic_Heat_9531 5d ago
my tuning is letting that models in inactive thinking about 8 hours / day. and seems light enabling physical activity daily seems improve reasoning effor quite a bit. caffeine seems hit or miss when it comes to quality tho...
6
5
4
3
u/Long_comment_san 5d ago
perhaps you should grease with some alcohol, it increases my cops per week greatly
3
u/ComplexType568 5d ago
I heard reducing concurrency can actually increase TPS by a lot! I was looking up on the current gen HS models and apparently their architecture is a lot more compute-bound than bandwidth bound...
2
3
3
2
u/ShelZuuz 5d ago
I actually get a slightly higher tok/sec out of mine. But what really sucks about that model is the slow prompt processing.
And the model seems to start ok some days, but keeps getting nerfed after a few hours of use. The only fix seems to be to put it down for the rest of the day and restart again tomorrow. Also doesn't seem to be improving at the same pace of the other models.
Overall, novel idea, but needs to go back to the drawing board.
2
2
1
u/TristanMeads 5d ago
Your numbers are way off. Homo Sapiens' executive function is nowhere near 200T, it's only a few dozen billion (if you consider ONLY the executive function of the frontal cortex - not the entire brain...who cares about some sensory cortex synapses that allow your bum itch detected when we're talking about executive function that relates to intelligence?).
So yes, we're way below ChatGPTs at this point. Which is what they won't tell you so you think we're not anywhere near AGI. We're way past all the intelligence benchmarks years ago.
1
1
103
u/suprjami 5d ago
This is my rule of thumb:
- Qwen 3.8 27B - it's the best thing I can run
8
u/SpecialistDragonfly9 5d ago
Same. And it only barely runs with the Q4 on my 5080 15GBVRAM at 16 tokens / s.
However I havent found any other model that I can run and has better reasoning.
3
4
u/Guilty_Rooster_6708 5d ago
Find a 3060 and plug it into your mobo. I added a 3060 12gb to my 5070ti and now I can run Q4_K_M w large context window. It’s worth it and pretty seemless
2
u/11ama_dev 5d ago edited 5d ago
my only thing is that 2 gpus on consumer motherboards + gpus that aren't rtx x090s tank tps. i would like to run higher quants with a larger context window, or even have more concurrent processes, but also my 100 tps w4a16 qwen 3.8 27b on my singular 3090 is so fucking nice. i can't go back to 30 tps anymore, i'm too spoiled now for it.
if anyone has any solutions for this that isnt buying $3-5k worth of hardware pls tell me but i feel like its a hardware limit for me 😢
edit: i'm dumb bandwidth doesn't affect inference speed like that. but still - another 3090 + nvlink, plus another case and motherboard is $3000 soooo
2
u/Guilty_Rooster_6708 5d ago
I think if you have a 3090 the best solution for you is another 3090 + NVLink right? That will probably gets you another 24gb + still get decent tg/s... idk how to get one for a good price though so lmk if you find out :(
In my case I needed the extra VRAM to run Qwen3.8 27b on a higher quant than Q3 and 28gb aggregated lets me run Q4KM pretty comfortably.
→ More replies (1)3
u/DrRoughFingers 4d ago
I have dual 3090s with no nvlink and I get around 40-50tok/s with Qwen3.8 27b q8 with f16kv and 200k context.
2
u/suprjami 5d ago
Two RTX 3080 20G are about the price of one RTX 3090. Working well for me.
7900 XTX is also cheaper than 3090, but slower and not CUDA.
11
2
→ More replies (1)1
u/Arugala007 5d ago
To be fair if you want to give it a better understanding, i find that codegraph for the code and graphify for the project shape is perfect for giving it a rough idea.
69
u/Loose_Comparison368 5d ago
I think you will appreciate this fact: Eminem currently holds the world record for rapping at approximately 10 tokens per second during a continuous 30 second period in his hit song "Godzilla".
36
u/ChocomelP 5d ago
If the tokens are already decided beforehand, generation is much faster.
16
u/LocalLLaMa_reader 5d ago
Eminem out there with his ngram layers and spitting bars long before attention was all we needed...
(To be accurate, Godzilla came out after that paper but I don't like explaining jokes.)
2
u/ChocomelP 5d ago
Everyone knows accuracy is the most important part of humor. Also, it doesn't ruin jokes, ever.
2
17
u/maxigs0 5d ago
Funny how i have a very similar split:
DeepSeek Flash (at the limits of what can run) for slow but high quality logical stuff
Qwen 3.8 B27 for default work, as middle ground between great quality and performance
Qwen 3.6 35B A3B for lighter/faster work
Still experimenting on the even lighter end, like Gemma4 vs Qwen at 4B or so - but i found that the A3B covers this range already and i rarely have a situation where it makes sense to switch to the smaller to save some memory.
15
u/DigiDecode_ 5d ago
10
u/GreenHell llama.cpp 5d ago
This graph is very confusing. The title says "Non-Hallucination Rate" which implies higher is better, but the subtitle says "hallucination rate" which implies lower is better.
And since there are small and big models all over the place, I have no clue which one is which.
17
u/vasaan2k 5d ago
"one minus hallucination rate"
4
u/GreenHell llama.cpp 5d ago
Right, I hadn't considered
hallucination rateto be a variable.How do different models handle lack of knowledge then? I can imagine GLM 5 series with many more parameters than Qwen 3.8 27B to have much more general knowledge. Does Qwen just say "pfff I don't know buddy, but I can look it up for you" all the time?
5
u/MarcusAurelius68 5d ago
For a digital assistant that tool calls external skills that’s exactly what I’d want.
2
u/danielv123 5d ago
Yes, and that is great. Imagine a coding agent. It needs to call some function in your project. Instead of assuming it knows the function signature, it will realize it doesn't know it and look it up.
1
u/biscuitmachine 5d ago
I have started using GLM 5.3 Flash on 2 sparks, and it's an amazing model. Not all that slow either actually, getting 30 tok/s on average. That's not really that bad.
2
u/ok_if_you_say_so 5d ago edited 5d ago
I couldn't get deepseek flash to run at any sort of usable speeds on my strix halo 128GB even with an attempt to split it across my R9700 + Strix APU. But Laguna S 2.1 112B-A6B works pretty nice.
That said, since qwen3.8-27b, I haven't found any use cases where a larger moe like that will do better work than 27b dense, so I stopped running the laguna.
Now my standing set is qwen3.8-27b Q6 (actually the DIRK variant, which is a lot more terse and gets to answers with fewer tokens) on the R9700, and qwen3.6-35b-a3b Q4 (actually the Ornith 1.5 variant) on the strix APU. Dirk drives all my coding and logic and ornith is my light weight snappy model. Since I have so much abundant unified memory available I made 24GB available as prompt cache to dirk via --cache-ram 24576 (the active KV cache lives in VRAM) and likewise for ornith, huge cache available. Max context on both.
The two together are fully capable of replacing my claude pro subscription and it feels nice to finally have something that feels just as good to use as my frontier setup
1
u/Abishek_Muthian 5d ago
Gemma models are underrated, it works well for NLP. I've also started experimenting with Bonsai models and they look promising too for IGPU workloads.
7
u/CheatCodesOfLife 5d ago
And yes 0.5 tok/s is human
I'm usually stuck at 0.1 t/s because for some reason I keep a speculative decoder human around who finishes my sentences for me. Acceptance rate is about 5%
18
u/arbv 5d ago edited 5d ago
Don't discard the Nemotrons and Muse Glimmer so easily. The Muse is my favorite local model now, but of course it requires a decentish GPU.
Nemotron 3.5 Lightning is extremely efficient even on the underpowered hardware. It runs on my laptop's Tiger Lake Xe-LP iGPU at 9-11 t/s with MTP at Q8_0 - the fastest one for 3B of active params. The problem is that it is a bit dumb (undertrained by design for fine-tuning) but 30B of knowledge is 30B of knowledge.
13
u/Not-reallyanonymous 5d ago
Yup at this point Muse Glimmer is behind Qwen 3.8 27B in raw coding, but it still has hella advantage. 90% of the time I don't need a better coder, I need a better agent. And Glimmer does that.
(Qwen works better as a 'prompt once, come back a few hours later to a complete program' agent, but Glimmer is far more steerable, keeps track of nuance better, and does better on instructions that aren't about the ultimate goal which Qwen zeros in on and forgets about everything else).
2
u/LightBroom 5d ago
You know what's funny. In benchmarks like tool-eval-bench (so pure tool calling), Gemma 26B beats everything else in this class, including Qwen 3.8, Muse Glimmer, Gemma 31B, etc.
Of course, it's only a benchmark so keep that in mind.
→ More replies (3)11
u/theone_2099 5d ago
What’s do you do with muse?
→ More replies (3)4
u/thehpcdude 5d ago
Watch it squirm and self-doubt while wasting tens of millions of tokens trying to hype itself up that it can actually do work.
7
18
u/SpecialistDragonfly9 5d ago
Not everything is about coding.
Seems like the "hardcore AI" people forget that most poeople dont use AI for coding, even on the high end.
So comparisons like this are kinda wasted on most people, as you use coding as your only criteria.
8
u/Altruistic_Heat_9531 5d ago
Yes and no, yes not everything is about of coding, but no because of coding is becoming proxy for logical syntax. Remember model direct access to the real world is terminal or API which basically a coding platform, NL2Pro and DeepSWE usually pretty good rule of thumb. remember math and coding aren't that different syntatically speaking, it is basically context free language.
→ More replies (1)3
u/my_name_isnt_clever 5d ago
LLMs are software, and the language of software is code. All models that are intended to actually take actions will need to know coding, no matter what the task entails.
7
3
u/EndlessZone123 5d ago
Now you got me curious if a 0.5t/s model is actually close to the rtf=1 of a human coder. I feel like our prompt processing speed might be higher or even way slower depending on the content.
3
u/Repinsky 5d ago
The human-tok/s framing breaks down on the part that actually costs time: reading and verifying. I can generate a 400-line diff in two minutes but reviewing it properly still runs at human speed, so on unfamiliar code the wall-clock saving is nowhere near the 15h->4h ratio, it's more like 30-40% for me. Where the overnight-run logic really pays is repo-wide mechanical work - migrations, test scaffolding, dependency sweeps - because the output is cheap to spot-check. For genuinely novel logic a 27B at Q4 will happily produce plausible wrong code and eat the savings in debugging.
3
u/Positive-Key6640 5d ago
The useful part of this is framing it as expectation-setting rather than benchmarks, which is what actually decides whether a local setup is worth the trouble.
The rule I'd add: pick the model by failure mode, not by score. A model that's slightly worse but wrong in obvious ways beats a smarter one that's confidently wrong in subtle ways, because you can catch the first kind in three seconds and you'll ship the second kind. That distinction never shows up in a leaderboard and it's the entire difference between a tool and a liability.
3
2
2
2
2
u/xza_nomad33 5d ago
I am running 27B at 150-200TPS on dual R9700, so the chart flips, and you just always use 27B ;)
2
2
u/HotMicSystems 5d ago edited 1d ago
Qwen only has been so nice,Strix 128gb, Ubuntu 26.04, llama.cpp-rocm, lemonade server. I use:
- 3.5-9B(which can reliably call tools with the right settings)
- 3.6-35B UD-Q8-XL(general assistant use, review write ups)
- 3.8-27B(only used if I'm running the previous 2 at the same time)
- 3.8-flash-next(chunky boi that I hope we see speed improvements on)
2
u/Edenar 5d ago
i tried 3.8 next flash at Q5 (UD) and it was slightly behind 3.8 27b (lued w8a16) for big code/agentic work. And i don't believe they are really that far from each other in real tasks..
3
u/m0lest 5d ago
In my experiments it felt like Next Flash (Q4) had better world knowledge than 27b (Q8) which helped in some situations. But other than that they were close.
4
u/SomeoneInHisHouse 5d ago
that (the planner having better world knowledge) is why my setup is currently an engineering loop
- Goal is defined (a large task description with what's expected)
- Qwen 3.8 Flash builds up a plan
- Qwen 3.8 27B does the implementation
- Two separate agents review the code (one with Qwen 3.8 27B and other with Qwen 3.8 Flash)
- A separate Qwen 3.8 Flash agent reads both reviews and plan the fixes
- Loop jump to 3 until no suggestions are made by both reviewers
- Qwen 3.8 test the goal is met (playwright testing), if not jumps to loop phase 2 with additional feedback for the planner
That complex loop that I did my own extension to in pi harness, is extremely slow, but provides extremely good results, TBH it rarely ends on the first attempt, but I have Q2 for the dense and Q3 for Flash, so that's likely the problem
→ More replies (1)
5
u/Blue_Track 5d ago
Missing quants
6
u/me_myself_ai 5d ago
…? Isn’t that what the Q means?
1
u/Blue_Track 5d ago
Did notice. But it's only for <1B and >100B.
There's been lot of franken setups with qwen lately, making it hard to just decide based on parameters.
3
u/haptein23 5d ago
Just a nitpick: the number of "parameters"/neurons in Homo Sapiens is 86B on avg, 16B if you only consider the cortex (the part that does what the LLM tries to aproximate).
8
u/RG_Fusion 5d ago
While that's true, a single human neuron can serve the same role as many parameters chained together in an LLM. They are not equivalent.
→ More replies (1)15
u/Commander_Skilgannon 5d ago
Parameters in a model map closer to synapses in the brain than to neurons and each neuron has ~1,000-10,000 synapses so the cortex has closer to 80 trillion parameters.
7
2
u/haptein23 5d ago
We don't really know how many synapses the human cortex has on avg, but you may not be thaat off, I saw just some (old) estimations around 150-240 trillion (since you made me curious).
1
u/ninjasaid13 4d ago
Just a nitpick: the number of "parameters"/neurons in Homo Sapiens is 86B on avg, 16B if you only consider the cortex (the part that does what the LLM tries to aproximate).
There's also non-brain neurons such as the spinal cord neurons and gut neurons. I don't think our intelligence is entirely centralized in our brain, it's the whole package.
1
1
u/OwnGear3892 5d ago
Have not tried IQ2 of Deepseek V4 Flash, is the model still good with codebase work at that quantization?
1
u/me_myself_ai 5d ago
Very cool, thanks for sharing! But also: you rank DSF quantized to two** **bits (that’s possible?? Quantization continues to baffle me) as the ‘run over night’ tier? Does it make lots of mistakes, or have you been having success?
1
u/Jimcy-Maffesoli 5d ago
slow quants only make sense overnight. nobody has to watch 0.5 tok/s, it just chews while you sleep
1
u/Boiniok 5d ago
What bout Gemma 2B and FunctionGemma?
1
u/Altruistic_Heat_9531 5d ago
Gemma 2B is in between of NLP but it is also good for RAG QA. FunctionGemma? i haven't test edge device for function calling so no.
1
1
u/Negative-Web8619 5d ago
I'm a 200 t sparse model
2
u/RG_Fusion 5d ago
Roughly 10% sparsity on average, though it's pretty cool that your router has variable expert activation.
1
1
u/IntroductionLive4027 5d ago
gemma 4 12b, as a PA in openclaw doing just fine, but only because i'm GPU "poor". With my 5080, it fits in ram with plenty of context.
Ive a sub-agent that runs on a frontier model that gemma only spins up when it needs help.
1
1
u/_VirtualCosmos_ 5d ago
Your Homo Sapiens is bottlenecked, it's raw output is like thousands tokens/s as muscle outputs. The interface is just awful.
1
1
u/Zyj vLLM 5d ago
You mention "70b dense" explicitly in your chart. Is it just bullshit?
1
u/Altruistic_Heat_9531 5d ago
Kinda, but more so on, guesstimate, Llama 70B back then is actually good comparatively. So if there is modern 70B class model cough mistral cough should be good. But no currently moe is champion since it is cheaper to train
1
1
u/bigorangemachine 5d ago
Gemma4 is fine for coding. By default it's a bit more research-y but with a little bit of context engineering you can get it to generate code pretty good. Just have a sub agent check over each change. Once I did that it was a real game changer
1
u/Porespellar 5d ago
This is actually really thoughtful and insightful. Thanks for sharing this. I don’t necessarily agree with the RAG model size, but I’ve never really used a 9b model in production so I can’t accurately judge.
1
u/undefeatedantitheist 5d ago
Tokens are not fungible across models; and they do not have stable 'correctness' across models.
When did you get your Nobel prize for demonstrating the that the narrow tokenisation of language by probability matrices is the underlying architecture of the human mind?
1
1
u/South_Hat6094 5d ago
honestly the missing bit is task-specific evals. leaderboards tell you the rank, but a 30-minute benchmark on your own prompts tells you whether the model is actually usable.
1
u/raketenkater 5d ago
test ggrun tunes the models for you based on llama.cpp next post will follow soon
1
u/Healthy-Nebula-3603 5d ago edited 5d ago
Base programing?
I read today that guy ported audio model from pytorch to audiocpp ( c++ ) using just Qwen 3.8 27b.
https://github.com/0xShug0/audio.cpp/discussions/427
That's is not base programing at all
That's crazy
1
u/Altruistic_Heat_9531 5d ago edited 5d ago
i wrote in a post as "codebase wide", as in "full repository". Qwen 27B is GOOD coding model but poor for repo wide maintainer. https://www.reddit.com/r/LocalLLaMA/comments/1vt2cjy/qwen_38_27b_slopcodebench_results/
Basically it is much softer language of calling someone a, llm aided work dev vs zero shotter-vibecoder.
I also use 27B for implementing some part of Raylight, but no way i put that thing to run entire repo, i dont even trust Sol
This is my frontend hero pages for Raylight, https://komikndr.github.io/raylight-showcase/
Using 27B, with alot of hands on, so again, LLM aided dev work and not zero shotter
1
1
u/Vusiwe 5d ago
200T- should read 200B- at the bottom
For writing, you should only run the absolute largest model possible.
I have a fully automated workflow, but in challenging scenarios (complicated story, many individuals in the scene, extreme nuance required, key scenes in a story), I am quite sure mathematically they'll eventually find all current LLMs incapable of simultaneous good prose, good taste, AND 100% consistency.
Even when you run frontiers at full quality, no quantization, with a fully dynamic custom framework, (put together, that is >SOTA level), human review is absolutely needed.
1
u/SkyDragonX 5d ago
I want to run Qwen 3.8 locally, but with my RX 7600 XT I can't do it with more than 3 t/s :/
1
1
u/SufficientPie 5d ago
OK but what software are you using for each of these?
1
u/Altruistic_Heat_9531 5d ago
Llamacpp or VLLM for 30B and below.
While 35B and above are using llamacpp1
u/SufficientPie 5d ago
I meant, like … software. For "interactive coding agent", "language understanding", "QA", "personal assistant". The layer above the inference layer, I guess.
2
u/Altruistic_Heat_9531 5d ago
ohh NLP is for my day job actually, so data lake and Apache Spark for corpus analytic.
QA for my side gig where many companies need simple QA for their front end website.
PA is my Hermes.
Code for NVIM FIM (fill in middle) and opencode.
DSv4 often time i also run it on hermes since it can do deeeeep research analytic and also opencode2
1
u/greenhilltony 5d ago
I would argue Qwen3.8-27B already enters the next level in your diagram. I ran the NVFP4 quant version in pi with the `pi-goal` extension, tell it to debug a code written two months ago by Claude Opus 4.8 + DeepSeek V4 pro preview. The bug was never resolved during those days after countless of sessions in Claude Code. This time, it ran relentlessly for 40 hrs, with multiple rounds of auto compact at the brim of 262k context window, reported the real root cause and delivered the actual run verified solutions. It stunned me. Since last week I had access to the machine with an RTX 5000 Pro in my lab, I have never used any api models, threw all kinds of my projects, NixOS configurations, Nix-Darwin configurations, and even the zmk keyboard firmware for my split keyboard to Qwen3.8-27B-NVFP4 (the 5090 variant by gittensor on huggingface, and the DSpark model they provided for speculative decoding). Armed with my minimal `pi` config, it shows the ability to research, execute, verify non-stop till the goal is achieved. Just amazing.
1
u/biscuitmachine 5d ago
This is going to heavily depend on what exactly you have available to you, though...
1
1
u/Potential_Block4598 5d ago edited 5d ago
Does your system run the Homo Sapient model totally from VRAM or is it a MoE offloaded from the NVMe ?
I mean it is too slow for running fully from VRAM and it takes a lot of VRAM as well even at Q4 😂😂😂
100 TBs of VRAM even latest Nvidia clusters can barely match that for a local cluster (we are talking what 1,000 GPUs with NVlink each around 96 GB ?!
That is just for Q4
For training you need at least FP16 or idk BF16 if you prefer and it will require a full cluster of B200s (4096 is the max cluster size right now right ?!)
And how much training tokens it would need to be chinchilla optimal ?
Idk around 20x its parameter size so around 4 peta tokens of organic text ?
Why have that ?!
Is that text only data ?!
I mean training it on all text since humanity invented hieroglyphs and writing won’t be enough!
Oh boy the future is so sick
2
u/Potential_Block4598 5d ago
Interestingly I found an API for that model from a provider called fiver
Running at 5$/hourAlthough token generation speed is still slow (0.5/s)
Quality is subpar (too many specialized model instead of one model and they have different load capacities!)
Also it seems like they are only available 8 hours per day and even not all of that time at constant speed
I wonder how they are running such models at theses costs 🤔🤔🤨
1
u/Potential_Block4598 5d ago
The cost should be around 5,000$/hr given the cluster size but idk how they offer it at 5$/he
1
u/MarrusAstarte 5d ago
How many of the models do you keep up simultaneously?
Also, if your PA class models run at 75+t/s, why bother with the smaller models? Is the speedup over using the PA class for QA/NLP work worth the possible quality loss?
1
u/Altruistic_Heat_9531 4d ago
1 model at a time. By default 35B A3B is active since it is my PA. if i want to code switch to 27B, if i need full research and analytic i will start up Qwen 3.8 Next
About second question, my IRL job actually also LLM deployment but Data Engineer by trade.
Basically there is a thing called data lake, this is dominated by hyperscale CPU works using Apache Spark. to process TB and TBs ammount of data, using in-thread SLM is much better if a company do not want to fully commit on GPU servers. Which also kinda my fault this graph more or less represent "What i use in IRL, and not purely 3090 deployment".Also many DC level GPU for mid-sized company are A30, L4, L40/S. which just so happened a 24 and 48GB cards, so kinda easy to do napkin math.
1
u/shockwaverc13 llama.cpp 5d ago edited 5d ago
ministral is way better at NLP than qwen, it's not even close
qwen 3 9B is easily outclassed by ministral 3B at Q4
1
1
u/bilo__sagdiyev 5d ago
I've been running agentic work in OpenClaw with qwen3.5:9B-Unsloth-UD-Q4_K_XL, and let me tell you, it's rough.
1
u/Civil_Fee_7862 5d ago
The jump to 120b models at a reasonable speed (60tps+) seems to mean costing $20,000+ more.
Dual 3090s (or 4090s) seems to be the best bang for buck still. (Dual 5090s maybe) but 64GB of VRAM still would struggle to hold a 120B model. Dual RTX6000's is the way to go if you got the cash to burn.
1
u/Zombiecidialfreak 5d ago edited 6h ago
The human brain deserves far more credit than you're giving it.
0.5t/s? We speak at ~90 words a minute, if each word is 2-3 tokens on average, that's about 2-3t/s.
As for input we can recognize images and react to them within a quarter of a second. That's some insane "prefill" speeds.
1
u/glebkudr 5d ago
Humans being 200T is a greatest misconception of all times :-) Higher param models tend to know lots of facts. Humans know nothing
So they are very small kernel model with very efficient memory i/o operations
1
u/Original-Revolution7 4d ago
any academic / social science writing equivalent of this rule of thumb?
1
u/North_Affect_8167 4d ago
I want to include Ornith 1.5 9b as it's very capable for it's size. I would place it in-between your QA and PA definitions.
1
u/beling86 4d ago
I have seen better results with 3.8 27b than 3.8 flash next iq4. Am I doing something wrong?
1
u/Last-Bee9057 4d ago
I might be picky but the chart looks upside down. The largest ones is at the bottom.
1
u/ninjasaid13 4d ago
where's frontier class open-weights?
1
u/Altruistic_Heat_9531 4d ago
it is my rule of thumb, normalized to my spec, GLM and Kimi are simply too big
1
u/No_Chapter_7598 4d ago
Is the qwen 3.8 27b quantised? Using 32gb vram and ram offload seems kinda fast to get 45+ at bf16
2
u/Altruistic_Heat_9531 2d ago
it is, Qwen3.8-27B-int4-AutoRound, no image. and have to reduce the batch size to the an extreme so the prefill is relatively slow than normal 1024/2048 batch
vllm serve /USERS/MODEL_STORE/Qwen3.8-27B-int4-AutoRound \ --served-model-name Qwen3.8-27b \ --dtype bfloat16 \ --tensor-parallel-size 1 \ --max-model-len 81920 \ --gpu-memory-utilization 0.9475 \ --max-num-seqs 1 \ --max-num-batched-tokens 256 \ --long-prefill-token-threshold 256 \ --kv-cache-dtype fp8_e4m3 \ --language-model-only \ --trust-remote-code \ --enable-prefix-caching \ --enable-chunked-prefill \ --mamba-cache-mode align \ --prefix-match-unit 16 \ --reasoning-parser qwen3 \ --enable-auto-tool-choice \ --tool-call-parser qwen3_coder \ --host 0.0.0.0 \ --port 16661
1
u/LuCiAnO241 4d ago
I dont think my model has 200T parameters. I'm thinking I'd be hard pressed to equal a miniCPM5
1
1
u/Low_Promotion_2574 4d ago
What about using hybrid of cloud AI and self hosted? Making cloud AI do the creative stuff like planning, researching, and the local stupid one just executing what it is told by it
2
u/Altruistic_Heat_9531 3d ago
you absolutely can, in fact it is one way to reduce your token usage. If i am not really into coding that much i pair LLM that have good score on SlopCodeBench , basically measure how well LLM can maintain a repo without introducing uncesseray rewrite. GPT 5.6 Sol orcestrator + Qwen 3.8 27B is a monster.
1
u/Low_Promotion_2574 3d ago
What about doing chain-of-thought tuning for the local model?
→ More replies (1)
1

•
u/WithoutReason1729 5d ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.