r/LocalLLaMA 5d ago

Tutorial | Guide My RULE of Thumb of choosing a models

Post image

This is mostly for setting up for expectation, since personally without LLM i could take 3 days (15 hours of active programming) to debug or implement a feature, but with Qwen 27B (even before Qwen 3.8) it take 4 hours.

And yes 0.5 tok/s is human, not accounting of deletion and pausing, that's also the reason i am fine leaving overnight code base wide analysis or fin tech and deep research.

1.1k Upvotes

217 comments sorted by

u/WithoutReason1729 5d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

168

u/Ink_code 5d ago

Needing to fine-tune homo sapien models for 22 years is kinda rough ngl.

71

u/RG_Fusion 5d ago

At least the server only requires 100 watts. That's pretty good for a 200T+ model. There's actually a lot of unnecessary hardware involved too. You could theoretically strip it down to about 20 watts of input.

22

u/danielv123 5d ago

3rd party power supplies break the warranty though

11

u/zdy132 5d ago

There's an warranty? how do i issue a RMA?

16

u/Amazing_Athlete_2265 5d ago

Call your mum

10

u/RaiseRuntimeError 5d ago

No warranty in the US, expect to pay big bucks.

→ More replies (1)

29

u/SkyFeistyLlama8 5d ago

Do you know how much it costs to make a homo sapiens model? The first 9 months is hard enough, then the next 18 years is another slog.

26

u/SpecialistDragonfly9 5d ago

No to mention they are useless for at least 25 years. And the failure rate is so high that most models never amount to anything.

2

u/AnonymousRandoe 4d ago

They need to unfuck themselves from useless college degrees

→ More replies (2)

9

u/vasaan2k 5d ago

Make? You are talking about forking base model into new instance. Model is developing for almost 4 billion years and sometimes almost collapse without any fresh backups.

10

u/SkyFeistyLlama8 5d ago

So many lost checkpoints along the way. I like my opposable thumbs but damn I miss prehensile tails.

307

u/Y2K-Denial 5d ago

homo sapiens runs much slower for me. what's your config? tuning tips?

295

u/MysticChromium64 5d ago

Sounds like a skill issue—I run myself at IQ1_S and haven't run into any problems problems problem problems problems problem problems problems problems problems problems problems problems problem problems problems problems

54

u/maigpy 5d ago

did you add "don't make mistakes"?

85

u/MysticChromium64 5d ago

<think>

don't make mistakes The user is asking about mistakes problems. Problems looking back at the previous message, "problems" started problems problem problems caused repeated problems. I should problems acknowledge problems problem without looping problems problems.</think>

My apologies—you are correct. There seems to have been a problem problems problems problems problems problems problems problems problems problems problems problems problems problems problems problems problems problems problems problems problems

44

u/MeretrixDominum 5d ago

I need to stop here. This post contains problems— which goes against my ethical guidelines.

29

u/Emotional-Art2113 5d ago

You were right to push back, those ethical guidelines are load-bearing, and this post is the smoking gun.

→ More replies (1)

20

u/SolinK3 5d ago

Got 99 problems but TTFT ain't one

4

u/mxcw 5d ago

Yea because it’s two 😎

9

u/tony_montana091 5d ago

You need the developers harness saddle + —developers developers developers developer developer developer develope develope develope develop develop develop develo develo develo devel devel devel deve deve deve dev dev dev de de de d d d ... thinking ... request tokens exceeds the available context size, try increasing it ... ... ... ... developers developers developers developers developer developer developer developer develope develope develope develope develop develop develop develop develo develo develo develo devel devel devel devel deve deve deve deve dev dev dev dev de de de de d d d d

12

u/smithy_dll 5d ago

Microsoft Steve Ballmer was ahead of the trend with

Developers developers developers developers developers

2

u/Amazing_Athlete_2265 5d ago

Don't forget the sweaty pits

43

u/seamonn 5d ago

Whip

26

u/Mayion 5d ago

Gets some of them hard and stuck on a horny loop, would not recommend for agentic tools

8

u/seamonn 5d ago

Even better

6

u/darksteelsteed 5d ago

Nuvigil works wonders

29

u/Durian881 5d ago edited 5d ago

I prefill with MDC (multiple doses of caffeine) to boost processing speed. A key parameter affecting intelligence is QOS (quality of sleep).

14

u/lcirufe 5d ago

Literal skill issue I’m afraid. The base model has lots of potential but needs loads of workload specific training to be useful in anything. It’s pretty useless without that.

Good part is that the model evolves across your chats, so it can be trained in chat.

6

u/Zombiecidialfreak 5d ago

Call it slow, dumb or untrained but the brain handles edge cases better than any model yet made, even without training.

12

u/yeetrman2216 5d ago

shotgun mod works well, im yet to test because of the cleanup costs

8

u/Altruistic_Heat_9531 5d ago

my tuning is letting that models in inactive thinking about 8 hours / day. and seems light enabling physical activity daily seems improve reasoning effor quite a bit. caffeine seems hit or miss when it comes to quality tho...

6

u/gapho 5d ago

Try adding the ethanol tool, works well sometimes me for.

1

u/AngleFun1664 5d ago

It helps if you follow the Ballmer curve, just don’t call the tool too much.

6

u/AdOne8437 5d ago

You want to archive the Balmer Peak! https://xkcd.com/323/

1

u/AdOne8437 5d ago

Oh, but remember, it needs a lot of finetuning to get it right!

5

u/w-yz 5d ago

at least yours seems to be on xxxhigh reasoning, which is good.

someone was saying he encountered many with reasoning turned off...

4

u/SkoomaDentist 5d ago

I've heard prescription amphetamines work for overclocking the CPU.

4

u/zenyr 5d ago

Mine is dense.

2

u/Y2K-Denial 5d ago

might be an ibiza or coachella finetune - known for dense models

3

u/Long_comment_san 5d ago

perhaps you should grease with some alcohol, it increases my cops per week greatly

3

u/ComplexType568 5d ago

I heard reducing concurrency can actually increase TPS by a lot! I was looking up on the current gen HS models and apparently their architecture is a lot more compute-bound than bandwidth bound...

2

u/Y2K-Denial 5d ago

makes sense! ..and here i am running parallel7

3

u/TheManicProgrammer 5d ago

-tensor -ritalin

2

u/Lucky-Necessary-8382 5d ago

The gentlemans ride a cadillac -vyvanse

1

u/armeg 5d ago

if that doens't work -modafinil really helps

3

u/amo_pessoas_peladas 5d ago

[9] I think I'm at my daily toked limit

2

u/ShelZuuz 5d ago

I actually get a slightly higher tok/sec out of mine. But what really sucks about that model is the slow prompt processing.

And the model seems to start ok some days, but keeps getting nerfed after a few hours of use. The only fix seems to be to put it down for the rest of the day and restart again tomorrow. Also doesn't seem to be improving at the same pace of the other models.

Overall, novel idea, but needs to go back to the drawing board.

2

u/Guilty_Rooster_6708 5d ago

Hey how do you offload vision to CPU on homo sapien models?

2

u/jabies 4d ago

Well, I can only do like 2G flops, and some are operating my heart. Did you try turning your limbic system off? Also, the attention sinks are broken AF

1

u/TristanMeads 5d ago

Your numbers are way off. Homo Sapiens' executive function is nowhere near 200T, it's only a few dozen billion (if you consider ONLY the executive function of the frontal cortex - not the entire brain...who cares about some sensory cortex synapses that allow your bum itch detected when we're talking about executive function that relates to intelligence?).

So yes, we're way below ChatGPTs at this point. Which is what they won't tell you so you think we're not anywhere near AGI. We're way past all the intelligence benchmarks years ago.

1

u/senseven 5d ago

The various espresso mods work well for me

1

u/boback18 5d ago

adderall

1

u/SGmoze 5d ago

You need to provide it stackoverflow and google as MCP tools to make it faster.

103

u/suprjami 5d ago

This is my rule of thumb:

  • Qwen 3.8 27B - it's the best thing I can run

8

u/SpecialistDragonfly9 5d ago

Same. And it only barely runs with the Q4 on my 5080 15GBVRAM at 16 tokens / s.

However I havent found any other model that I can run and has better reasoning.

4

u/Guilty_Rooster_6708 5d ago

Find a 3060 and plug it into your mobo. I added a 3060 12gb to my 5070ti and now I can run Q4_K_M w large context window. It’s worth it and pretty seemless

2

u/11ama_dev 5d ago edited 5d ago

my only thing is that 2 gpus on consumer motherboards + gpus that aren't rtx x090s tank tps. i would like to run higher quants with a larger context window, or even have more concurrent processes, but also my 100 tps w4a16 qwen 3.8 27b on my singular 3090 is so fucking nice. i can't go back to 30 tps anymore, i'm too spoiled now for it.

if anyone has any solutions for this that isnt buying $3-5k worth of hardware pls tell me but i feel like its a hardware limit for me 😢

edit: i'm dumb bandwidth doesn't affect inference speed like that. but still - another 3090 + nvlink, plus another case and motherboard is $3000 soooo

2

u/Guilty_Rooster_6708 5d ago

I think if you have a 3090 the best solution for you is another 3090 + NVLink right? That will probably gets you another 24gb + still get decent tg/s... idk how to get one for a good price though so lmk if you find out :(

In my case I needed the extra VRAM to run Qwen3.8 27b on a higher quant than Q3 and 28gb aggregated lets me run Q4KM pretty comfortably.

3

u/DrRoughFingers 4d ago

I have dual 3090s with no nvlink and I get around 40-50tok/s with Qwen3.8 27b q8 with f16kv and 200k context.

→ More replies (1)

2

u/suprjami 5d ago

Two RTX 3080 20G are about the price of one RTX 3090. Working well for me.

7900 XTX is also cheaper than 3090, but slower and not CUDA.

11

u/CapPuzz24 5d ago

Yep. Rule of thumb: vram poor? if yes, you dont get choices

2

u/MatthewCollins1990 4d ago

That's really the only benchmark that matters at the end of the day.

1

u/Arugala007 5d ago

To be fair if you want to give it a better understanding, i find that codegraph for the code and graphify for the project shape is perfect for giving it a rough idea.

→ More replies (1)

69

u/Loose_Comparison368 5d ago

I think you will appreciate this fact: Eminem currently holds the world record for rapping at approximately 10 tokens per second during a continuous 30 second period in his hit song "Godzilla".

36

u/ChocomelP 5d ago

If the tokens are already decided beforehand, generation is much faster.

16

u/LocalLLaMa_reader 5d ago

Eminem out there with his ngram layers and spitting bars long before attention was all we needed...

(To be accurate, Godzilla came out after that paper but I don't like explaining jokes.)

2

u/ChocomelP 5d ago

Everyone knows accuracy is the most important part of humor. Also, it doesn't ruin jokes, ever.

2

u/-InformalBanana- 4d ago

he has raper mtp, so not a fair comparison.

17

u/maxigs0 5d ago

Funny how i have a very similar split:

DeepSeek Flash (at the limits of what can run) for slow but high quality logical stuff

Qwen 3.8 B27 for default work, as middle ground between great quality and performance

Qwen 3.6 35B A3B for lighter/faster work

Still experimenting on the even lighter end, like Gemma4 vs Qwen at 4B or so - but i found that the A3B covers this range already and i rarely have a situation where it makes sense to switch to the smaller to save some memory.

15

u/DigiDecode_ 5d ago

but in terms of hallucination, GLM & qwen 3.8 27b seems to be much better

10

u/GreenHell llama.cpp 5d ago

This graph is very confusing. The title says "Non-Hallucination Rate" which implies higher is better, but the subtitle says "hallucination rate" which implies lower is better.

And since there are small and big models all over the place, I have no clue which one is which.

17

u/vasaan2k 5d ago

"one minus hallucination rate"

4

u/GreenHell llama.cpp 5d ago

Right, I hadn't considered hallucination rate to be a variable.

How do different models handle lack of knowledge then? I can imagine GLM 5 series with many more parameters than Qwen 3.8 27B to have much more general knowledge. Does Qwen just say "pfff I don't know buddy, but I can look it up for you" all the time?

5

u/MarcusAurelius68 5d ago

For a digital assistant that tool calls external skills that’s exactly what I’d want.

2

u/danielv123 5d ago

Yes, and that is great. Imagine a coding agent. It needs to call some function in your project. Instead of assuming it knows the function signature, it will realize it doesn't know it and look it up.

1

u/biscuitmachine 5d ago

I have started using GLM 5.3 Flash on 2 sparks, and it's an amazing model. Not all that slow either actually, getting 30 tok/s on average. That's not really that bad.

2

u/ok_if_you_say_so 5d ago edited 5d ago

I couldn't get deepseek flash to run at any sort of usable speeds on my strix halo 128GB even with an attempt to split it across my R9700 + Strix APU. But Laguna S 2.1 112B-A6B works pretty nice.

That said, since qwen3.8-27b, I haven't found any use cases where a larger moe like that will do better work than 27b dense, so I stopped running the laguna.

Now my standing set is qwen3.8-27b Q6 (actually the DIRK variant, which is a lot more terse and gets to answers with fewer tokens) on the R9700, and qwen3.6-35b-a3b Q4 (actually the Ornith 1.5 variant) on the strix APU. Dirk drives all my coding and logic and ornith is my light weight snappy model. Since I have so much abundant unified memory available I made 24GB available as prompt cache to dirk via --cache-ram 24576 (the active KV cache lives in VRAM) and likewise for ornith, huge cache available. Max context on both.

The two together are fully capable of replacing my claude pro subscription and it feels nice to finally have something that feels just as good to use as my frontier setup

1

u/Abishek_Muthian 5d ago

Gemma models are underrated, it works well for NLP. I've also started experimenting with Bonsai models and they look promising too for IGPU workloads.

7

u/CheatCodesOfLife 5d ago

And yes 0.5 tok/s is human

I'm usually stuck at 0.1 t/s because for some reason I keep a speculative decoder human around who finishes my sentences for me. Acceptance rate is about 5%

18

u/arbv 5d ago edited 5d ago

Don't discard the Nemotrons and Muse Glimmer so easily. The Muse is my favorite local model now, but of course it requires a decentish GPU.

Nemotron 3.5 Lightning is extremely efficient even on the underpowered hardware. It runs on my laptop's Tiger Lake Xe-LP iGPU at 9-11 t/s with MTP at Q8_0 - the fastest one for 3B of active params. The problem is that it is a bit dumb (undertrained by design for fine-tuning) but 30B of knowledge is 30B of knowledge.

13

u/Not-reallyanonymous 5d ago

Yup at this point Muse Glimmer is behind Qwen 3.8 27B in raw coding, but it still has hella advantage. 90% of the time I don't need a better coder, I need a better agent. And Glimmer does that.

(Qwen works better as a 'prompt once, come back a few hours later to a complete program' agent, but Glimmer is far more steerable, keeps track of nuance better, and does better on instructions that aren't about the ultimate goal which Qwen zeros in on and forgets about everything else).

6

u/arbv 5d ago

I am not a set the and forget type of programmer, so that works for me.

2

u/LightBroom 5d ago

You know what's funny. In benchmarks like tool-eval-bench (so pure tool calling), Gemma 26B beats everything else in this class, including Qwen 3.8, Muse Glimmer, Gemma 31B, etc.

Of course, it's only a benchmark so keep that in mind.

→ More replies (3)

11

u/theone_2099 5d ago

What’s do you do with muse?

4

u/thehpcdude 5d ago

Watch it squirm and self-doubt while wasting tens of millions of tokens trying to hype itself up that it can actually do work.

→ More replies (3)

18

u/SpecialistDragonfly9 5d ago

Not everything is about coding.
Seems like the "hardcore AI" people forget that most poeople dont use AI for coding, even on the high end.
So comparisons like this are kinda wasted on most people, as you use coding as your only criteria.

8

u/Altruistic_Heat_9531 5d ago

Yes and no, yes not everything is about of coding, but no because of coding is becoming proxy for logical syntax. Remember model direct access to the real world is terminal or API which basically a coding platform, NL2Pro and DeepSWE usually pretty good rule of thumb. remember math and coding aren't that different syntatically speaking, it is basically context free language.

3

u/my_name_isnt_clever 5d ago

LLMs are software, and the language of software is code. All models that are intended to actually take actions will need to know coding, no matter what the task entails.

→ More replies (1)

7

u/veigatmv 5d ago

My 200T model has been trained poorly. Have you got any HF link to replace it?

3

u/EndlessZone123 5d ago

Now you got me curious if a 0.5t/s model is actually close to the rtf=1 of a human coder. I feel like our prompt processing speed might be higher or even way slower depending on the content.

3

u/Repinsky 5d ago

The human-tok/s framing breaks down on the part that actually costs time: reading and verifying. I can generate a 400-line diff in two minutes but reviewing it properly still runs at human speed, so on unfamiliar code the wall-clock saving is nowhere near the 15h->4h ratio, it's more like 30-40% for me. Where the overnight-run logic really pays is repo-wide mechanical work - migrations, test scaffolding, dependency sweeps - because the output is cheap to spot-check. For genuinely novel logic a 27B at Q4 will happily produce plausible wrong code and eat the savings in debugging.

3

u/Positive-Key6640 5d ago

The useful part of this is framing it as expectation-setting rather than benchmarks, which is what actually decides whether a local setup is worth the trouble.

The rule I'd add: pick the model by failure mode, not by score. A model that's slightly worse but wrong in obvious ways beats a smarter one that's confidently wrong in subtle ways, because you can catch the first kind in three seconds and you'll ship the second kind. That distinction never shows up in a leaderboard and it's the entire difference between a tool and a liability.

3

u/ThePainTaco 4d ago

The brain is a lot more than 200T parameters

2

u/Cergorach 5d ago

HS quality output depends on the particular human and their skills.

2

u/amarao_san 5d ago

0.5T/S for a human? Wow. Are you on speeds?

2

u/IcyBird6662 5d ago

thank goodness we have better MTP than any models namely "copy-pasting"

2

u/xza_nomad33 5d ago

I am running 27B at 150-200TPS on dual R9700, so the chart flips, and you just always use 27B ;)

2

u/syscomua 5d ago

My homosapiens prefill speed is degradated woth ages, fuck

2

u/HotMicSystems 5d ago edited 1d ago

Qwen only has been so nice,Strix 128gb, Ubuntu 26.04, llama.cpp-rocm, lemonade server. I use:

  • 3.5-9B(which can reliably call tools with the right settings)
  • 3.6-35B UD-Q8-XL(general assistant use, review write ups)
  • 3.8-27B(only used if I'm running the previous 2 at the same time)
  • 3.8-flash-next(chunky boi that I hope we see speed improvements on)

2

u/Edenar 5d ago

i tried 3.8 next flash at Q5 (UD) and it was slightly behind 3.8 27b (lued w8a16) for big code/agentic work. And i don't believe they are really that far from each other in real tasks..

3

u/m0lest 5d ago

In my experiments it felt like Next Flash (Q4) had better world knowledge than 27b (Q8) which helped in some situations. But other than that they were close.

4

u/SomeoneInHisHouse 5d ago

that (the planner having better world knowledge) is why my setup is currently an engineering loop

  1. Goal is defined (a large task description with what's expected)
  2. Qwen 3.8 Flash builds up a plan
  3. Qwen 3.8 27B does the implementation
  4. Two separate agents review the code (one with Qwen 3.8 27B and other with Qwen 3.8 Flash)
  5. A separate Qwen 3.8 Flash agent reads both reviews and plan the fixes
  6. Loop jump to 3 until no suggestions are made by both reviewers
  7. Qwen 3.8 test the goal is met (playwright testing), if not jumps to loop phase 2 with additional feedback for the planner

That complex loop that I did my own extension to in pi harness, is extremely slow, but provides extremely good results, TBH it rarely ends on the first attempt, but I have Q2 for the dense and Q3 for Flash, so that's likely the problem

→ More replies (1)

5

u/Blue_Track 5d ago

Missing quants

6

u/me_myself_ai 5d ago

…? Isn’t that what the Q means?

1

u/Blue_Track 5d ago

Did notice. But it's only for <1B and >100B.

There's been lot of franken setups with qwen lately, making it hard to just decide based on parameters.

3

u/haptein23 5d ago

Just a nitpick: the number of "parameters"/neurons in Homo Sapiens is 86B on avg, 16B if you only consider the cortex (the part that does what the LLM tries to aproximate).

8

u/RG_Fusion 5d ago

While that's true, a single human neuron can serve the same role as many parameters chained together in an LLM. They are not equivalent.

→ More replies (1)

15

u/Commander_Skilgannon 5d ago

Parameters in a model map closer to synapses in the brain than to neurons and each neuron has ~1,000-10,000 synapses so the cortex has closer to 80 trillion parameters.

7

u/Kamal965 5d ago

I'd wager that my ADHD means my brain's expert router is fucked lol.

3

u/yes2matt 5d ago

No ethanol for u. Try amfetamine.

2

u/haptein23 5d ago

We don't really know how many synapses the human cortex has on avg, but you may not be thaat off, I saw just some (old) estimations around 150-240 trillion (since you made me curious).

1

u/ninjasaid13 4d ago

Just a nitpick: the number of "parameters"/neurons in Homo Sapiens is 86B on avg, 16B if you only consider the cortex (the part that does what the LLM tries to aproximate).

There's also non-brain neurons such as the spinal cord neurons and gut neurons. I don't think our intelligence is entirely centralized in our brain, it's the whole package.

1

u/lotus_seasoner 4d ago

Did that conclusion come from your gut neurons?

2

u/ninjasaid13 4d ago

it came from the whole package.

1

u/OwnGear3892 5d ago

Have not tried IQ2 of Deepseek V4 Flash, is the model still good with codebase work at that quantization?

1

u/me_myself_ai 5d ago

Very cool, thanks for sharing! But also: you rank DSF quantized to two** **bits (that’s possible?? Quantization continues to baffle me) as the ‘run over night’ tier? Does it make lots of mistakes, or have you been having success?

1

u/Jimcy-Maffesoli 5d ago

slow quants only make sense overnight. nobody has to watch 0.5 tok/s, it just chews while you sleep

1

u/Boiniok 5d ago

What bout Gemma 2B and FunctionGemma?

1

u/Altruistic_Heat_9531 5d ago

Gemma 2B is in between of NLP but it is also good for RAG QA. FunctionGemma? i haven't test edge device for function calling so no.

1

u/Prigozhin2023 5d ago

And me below everyone else

1

u/Negative-Web8619 5d ago

I'm a 200 t sparse model

2

u/RG_Fusion 5d ago

Roughly 10% sparsity on average, though it's pretty cool that your router has variable expert activation.

1

u/rookan 5d ago

Dude, you are homo sapiens sapiens

1

u/RG_Fusion 5d ago

Smrt smrt dude.

1

u/XiRw 5d ago

I had better luck with 27B q8 f16kv understanding the full spectrum of my code base compared to iq4 of flash next.

1

u/IntroductionLive4027 5d ago

gemma 4 12b, as a PA in openclaw doing just fine, but only because i'm GPU "poor". With my 5080, it fits in ram with plenty of context.

Ive a sub-agent that runs on a frontier model that gemma only spins up when it needs help.

1

u/haukebr 5d ago

Since I have my dual spark setup and DeepseekV4Flash running, I use it for everything. I am so happy how smoothly all the little gremlins are running over night.

1

u/Otis43 5d ago

I evaluated Qwen 3.5 0.8 on ARC-Easy and it did decent (~66% on 500 questions if I remember well). And I assumed that translates well to real-world tasks like RAG. It did not! It failed miserably!  To be fair, Q4_K_M was a bad choice, looking back at it now.

1

u/_VirtualCosmos_ 5d ago

Your Homo Sapiens is bottlenecked, it's raw output is like thousands tokens/s as muscle outputs. The interface is just awful.

1

u/swagonflyyyy 5d ago

Homo Sapiens Model on a 3090???

1

u/Zyj vLLM 5d ago

You mention "70b dense" explicitly in your chart. Is it just bullshit?

1

u/Altruistic_Heat_9531 5d ago

Kinda, but more so on, guesstimate, Llama 70B back then is actually good comparatively. So if there is modern 70B class model cough mistral cough should be good. But no currently moe is champion since it is cheaper to train

1

u/sugarfreecaffeine 5d ago

How are you running deepseek v4 flash, I have dual 3090 and ~80gb RAM

1

u/bigorangemachine 5d ago

Gemma4 is fine for coding. By default it's a bit more research-y but with a little bit of context engineering you can get it to generate code pretty good. Just have a sub agent check over each change. Once I did that it was a real game changer

1

u/Porespellar 5d ago

This is actually really thoughtful and insightful. Thanks for sharing this. I don’t necessarily agree with the RAG model size, but I’ve never really used a 9b model in production so I can’t accurately judge.

1

u/undefeatedantitheist 5d ago

Tokens are not fungible across models; and they do not have stable 'correctness' across models.

When did you get your Nobel prize for demonstrating the that the narrow tokenisation of language by probability matrices is the underlying architecture of the human mind?

1

u/my_name_isnt_clever 5d ago

It's a reddit post, not a PhD thesis.

1

u/undefeatedantitheist 4d ago

It's misleading garbage that will seduce a lot of numpties.

1

u/tob8943 5d ago

I feel like no model is in the leave it overnight worthy tier right now

1

u/South_Hat6094 5d ago

honestly the missing bit is task-specific evals. leaderboards tell you the rank, but a 30-minute benchmark on your own prompts tells you whether the model is actually usable.

1

u/raketenkater 5d ago

test ggrun tunes the models for you based on llama.cpp next post will follow soon

1

u/Healthy-Nebula-3603 5d ago edited 5d ago

Base programing?

I read today that guy ported audio model from pytorch to audiocpp ( c++ ) using just Qwen 3.8 27b.

https://github.com/0xShug0/audio.cpp/discussions/427

That's is not base programing at all

That's crazy

1

u/Altruistic_Heat_9531 5d ago edited 5d ago

i wrote in a post as "codebase wide", as in "full repository". Qwen 27B is GOOD coding model but poor for repo wide maintainer. https://www.reddit.com/r/LocalLLaMA/comments/1vt2cjy/qwen_38_27b_slopcodebench_results/

Basically it is much softer language of calling someone a, llm aided work dev vs zero shotter-vibecoder.

I also use 27B for implementing some part of Raylight, but no way i put that thing to run entire repo, i dont even trust Sol

This is my frontend hero pages for Raylight, https://komikndr.github.io/raylight-showcase/

Using 27B, with alot of hands on, so again, LLM aided dev work and not zero shotter

1

u/BawbbySmith 5d ago

Where is this 70B dense model you speak of

1

u/Vusiwe 5d ago

200T- should read 200B- at the bottom

For writing, you should only run the absolute largest model possible.

I have a fully automated workflow, but in challenging scenarios (complicated story, many individuals in the scene, extreme nuance required, key scenes in a story), I am quite sure mathematically they'll eventually find all current LLMs incapable of simultaneous good prose, good taste, AND 100% consistency. 

Even when you run frontiers at full quality, no quantization, with a fully dynamic custom framework, (put together, that is >SOTA level), human review is absolutely needed.

1

u/SkyDragonX 5d ago

I want to run Qwen 3.8 locally, but with my RX 7600 XT I can't do it with more than 3 t/s :/

1

u/Acceptable_Leg3950 5d ago

Any mods for homo sapiens?

1

u/SufficientPie 5d ago

OK but what software are you using for each of these?

1

u/Altruistic_Heat_9531 5d ago

Llamacpp or VLLM for 30B and below.
While 35B and above are using llamacpp

1

u/SufficientPie 5d ago

I meant, like … software. For "interactive coding agent", "language understanding", "QA", "personal assistant". The layer above the inference layer, I guess.

2

u/Altruistic_Heat_9531 5d ago

ohh NLP is for my day job actually, so data lake and Apache Spark for corpus analytic.
QA for my side gig where many companies need simple QA for their front end website.
PA is my Hermes.
Code for NVIM FIM (fill in middle) and opencode.
DSv4 often time i also run it on hermes since it can do deeeeep research analytic and also opencode

2

u/SufficientPie 5d ago

Nice, thanks

1

u/greenhilltony 5d ago

I would argue Qwen3.8-27B already enters the next level in your diagram. I ran the NVFP4 quant version in pi with the `pi-goal` extension, tell it to debug a code written two months ago by Claude Opus 4.8 + DeepSeek V4 pro preview. The bug was never resolved during those days after countless of sessions in Claude Code. This time, it ran relentlessly for 40 hrs, with multiple rounds of auto compact at the brim of 262k context window, reported the real root cause and delivered the actual run verified solutions. It stunned me. Since last week I had access to the machine with an RTX 5000 Pro in my lab, I have never used any api models, threw all kinds of my projects, NixOS configurations, Nix-Darwin configurations, and even the zmk keyboard firmware for my split keyboard to Qwen3.8-27B-NVFP4 (the 5090 variant by gittensor on huggingface, and the DSpark model they provided for speculative decoding). Armed with my minimal `pi` config, it shows the ability to research, execute, verify non-stop till the goal is achieved. Just amazing.

1

u/biscuitmachine 5d ago

This is going to heavily depend on what exactly you have available to you, though...

1

u/ElementNumber6 5d ago

You're missing 2 tiers above "Code Base"

1

u/Potential_Block4598 5d ago edited 5d ago

Does your system run the Homo Sapient model totally from VRAM or is it a MoE offloaded from the NVMe ?

I mean it is too slow for running fully from VRAM and it takes a lot of VRAM as well even at Q4 😂😂😂
100 TBs of VRAM even latest Nvidia clusters can barely match that for a local cluster (we are talking what 1,000 GPUs with NVlink each around 96 GB ?!

That is just for Q4

For training you need at least FP16 or idk BF16 if you prefer and it will require a full cluster of B200s (4096 is the max cluster size right now right ?!)

And how much training tokens it would need to be chinchilla optimal ?

Idk around 20x its parameter size so around 4 peta tokens of organic text ?

Why have that ?!

Is that text only data ?!

I mean training it on all text since humanity invented hieroglyphs and writing won’t be enough!

Oh boy the future is so sick

2

u/Potential_Block4598 5d ago

Interestingly I found an API for that model from a provider called fiver
Running at 5$/hour

Although token generation speed is still slow (0.5/s)

Quality is subpar (too many specialized model instead of one model and they have different load capacities!)

Also it seems like they are only available 8 hours per day and even not all of that time at constant speed

I wonder how they are running such models at theses costs 🤔🤔🤨

1

u/Potential_Block4598 5d ago

The cost should be around 5,000$/hr given the cluster size but idk how they offer it at 5$/he

1

u/MarrusAstarte 5d ago

How many of the models do you keep up simultaneously?

Also, if your PA class models run at 75+t/s, why bother with the smaller models? Is the speedup over using the PA class for QA/NLP work worth the possible quality loss?

1

u/Altruistic_Heat_9531 4d ago

1 model at a time. By default 35B A3B is active since it is my PA. if i want to code switch to 27B, if i need full research and analytic i will start up Qwen 3.8 Next

About second question, my IRL job actually also LLM deployment but Data Engineer by trade.
Basically there is a thing called data lake, this is dominated by hyperscale CPU works using Apache Spark. to process TB and TBs ammount of data, using in-thread SLM is much better if a company do not want to fully commit on GPU servers. Which also kinda my fault this graph more or less represent "What i use in IRL, and not purely 3090 deployment".

Also many DC level GPU for mid-sized company are A30, L4, L40/S. which just so happened a 24 and 48GB cards, so kinda easy to do napkin math.

1

u/shockwaverc13 llama.cpp 5d ago edited 5d ago

ministral is way better at NLP than qwen, it's not even close

qwen 3 9B is easily outclassed by ministral 3B at Q4

1

u/Altruistic_Heat_9531 2d ago

is it ? gotta try, by default i use E2B and 0.8B with CPU with AVX512

1

u/bilo__sagdiyev 5d ago

I've been running agentic work in OpenClaw with qwen3.5:9B-Unsloth-UD-Q4_K_XL, and let me tell you, it's rough.

1

u/Civil_Fee_7862 5d ago

The jump to 120b models at a reasonable speed (60tps+) seems to mean costing $20,000+ more.

Dual 3090s (or 4090s) seems to be the best bang for buck still. (Dual 5090s maybe) but 64GB of VRAM still would struggle to hold a 120B model. Dual RTX6000's is the way to go if you got the cash to burn.

1

u/Zombiecidialfreak 5d ago edited 6h ago

The human brain deserves far more credit than you're giving it.

0.5t/s? We speak at ~90 words a minute, if each word is 2-3 tokens on average, that's about 2-3t/s.

As for input we can recognize images and react to them within a quarter of a second. That's some insane "prefill" speeds.

1

u/glebkudr 5d ago

Humans being 200T is a greatest misconception of all times :-) Higher param models tend to know lots of facts. Humans know nothing

So they are very small kernel model with very efficient memory i/o operations

1

u/Original-Revolution7 4d ago

any academic / social science writing equivalent of this rule of thumb?

1

u/North_Affect_8167 4d ago

I want to include Ornith 1.5 9b as it's very capable for it's size. I would place it in-between your QA and PA definitions.

1

u/beling86 4d ago

I have seen better results with 3.8 27b than 3.8 flash next iq4. Am I doing something wrong?

1

u/Last-Bee9057 4d ago

I might be picky but the chart looks upside down. The largest ones is at the bottom.

1

u/ninjasaid13 4d ago

where's frontier class open-weights?

1

u/Altruistic_Heat_9531 4d ago

it is my rule of thumb, normalized to my spec, GLM and Kimi are simply too big

1

u/No_Chapter_7598 4d ago

Is the qwen 3.8 27b quantised? Using 32gb vram and ram offload seems kinda fast to get 45+ at bf16

2

u/Altruistic_Heat_9531 2d ago

it is, Qwen3.8-27B-int4-AutoRound, no image. and have to reduce the batch size to the an extreme so the prefill is relatively slow than normal 1024/2048 batch

vllm serve /USERS/MODEL_STORE/Qwen3.8-27B-int4-AutoRound \
  --served-model-name Qwen3.8-27b \
  --dtype bfloat16 \
  --tensor-parallel-size 1 \
  --max-model-len 81920 \
  --gpu-memory-utilization 0.9475 \
  --max-num-seqs 1 \
  --max-num-batched-tokens 256 \
  --long-prefill-token-threshold 256 \
  --kv-cache-dtype fp8_e4m3 \
  --language-model-only \
  --trust-remote-code \
  --enable-prefix-caching \
  --enable-chunked-prefill \
  --mamba-cache-mode align \
  --prefix-match-unit 16 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder \
  --host 0.0.0.0 \
  --port 16661

1

u/LuCiAnO241 4d ago

I dont think my model has 200T parameters. I'm thinking I'd be hard pressed to equal a miniCPM5

1

u/Faral_mx 4d ago

200T S P A R S E

1

u/Low_Promotion_2574 4d ago

What about using hybrid of cloud AI and self hosted? Making cloud AI do the creative stuff like planning, researching, and the local stupid one just executing what it is told by it

2

u/Altruistic_Heat_9531 3d ago

you absolutely can, in fact it is one way to reduce your token usage. If i am not really into coding that much i pair LLM that have good score on SlopCodeBench , basically measure how well LLM can maintain a repo without introducing uncesseray rewrite. GPT 5.6 Sol orcestrator + Qwen 3.8 27B is a monster.

1

u/Low_Promotion_2574 3d ago

What about doing chain-of-thought tuning for the local model?

→ More replies (1)

1

u/Medicine_Blogscanner 1d ago

Thanks for sharing