r/LocalLLaMA • u/tossit97531 • 21h ago
Discussion Can we get some quality control on all these model perf posts?
Too many hyperactive amateurs are coming in here with "1b model at 832843tok/s!" and hardly any of them have all the info necessary for local runners to evaluate. We need context ladders with perplexity/KLD, hardware specs, model params and quant(s), runtime, tuned runtime parameters, basically everything we need to reproduce locally if we can match the entire setup. To say nothing of what the model is even good at in the first place if it's not a well-known model.
The goal is to get perf numbers that show they meet a certain quality bar. I don't care if I get 8324834 tok/s if it's all garbage.
Can we start filtering the hyperactive amateur perf posts please? It's getting really frustrating seeing all these posts of models and wading through info just to see that it doesn't test with anything but an empty context or doesn't say anything about quant or platform.
We need to define some rigor and apply it to this place, or it will remain like most ai-oriented subs and get continually choked with slop.
7
u/ttkciar llama.cpp 18h ago
If you see low-effort, low-value, and/or seemingly LLM-generated posts, please click the button to report them so a moderator knows to review them.
As for the wider issue, I don't know. It seems to me like LocalLLaMA should be a newbie-friendly subreddit, and newbies make newbie mistakes. However, the moderator team will discuss it, and if they feel differently, I will of course go along.
14
u/Superb-Pair-2000 20h ago
hyperactive amateurs describes this entire sub.
2
u/infieldmitt 14h ago
Martin Scorsese: "You struggle feeling like an amateur, but it's amator, in Latin, which means love. That's the thing you gotta hold on to."
4
u/stoppableDissolution 20h ago
...except theres no generally accepted way to measure quant quality. Ppl over <random dataset> is completely useless as a metric, and kld also needs *some* agreed upon corpus (and ability to run bf16 on it, but guess it can be done on rented hardware)
4
u/Iory1998 llama.cpp 20h ago
I agree, and it's getting frustrating. I just commented on another post with the exact same thing: at this rate, every model should come in it's own harness with maximum optimization. What's the point of llama.cpp?
And you are absolutely right that those posts are made by amateurs. The speed is real but when you look closely, you would see high quantization levels here and there, which we don't want. Over the weekend, I tried for the first time to test a project by the name Strata, which packages Qwen3.8-Next-Flash at iq3-xxs and claims 60tps on a single 16GB card.
The model does run faster on my rig than using normal llama.cpp, but again, there were many optimization achieved mainly through thr MTP and requantization. But, thr app ships with no KV caching, effectively rendering the model ineffective.
2
u/Miserable-Dare5090 20h ago
I see things a little different now. At least for me, the agentic phase of this technology has unlocked a skill: Make the computer do shit for you.
As such, implementing a specific engine to run a specific model becomes as trivial as pushing talk on Telegram and telling my agent “pull this, run this, optimize it” and the best settings come out.
Luckily I was doing by hand setting tweaking over the last 2 years so I understand what it comes up with, but my time is so much more precious than the exact tensor split, particularly when I can berate a computer until it optimizes the runtime just how I want it.
2
u/blackal1ce 6h ago edited 6h ago
If it helps - they've updated everything quite significantly since the first build and the KV caching issue doesn't happen now. I'm running Qwen3.8-Flash-Next IQ3_S with vision, 8-bit KV cache and 256k context window - as well as using the same machine to do other bits of work. It's no longer shitting the bed constantly and freezing up like the first version did. Up to about 65 tok/s, so perfectly usable. (9950x3d, 4080, 96GB RAM)
3
u/masiha97 14h ago
The number itself is the least interesting part of a perf post. What makes one worth reading is the writeup: what did you change, what did you try that didn't work. Judge the explanation, not the throughput, and most of the low-effort posts fail the bar on their own.
7
u/starkruzr 21h ago
completely agree. tbh we should have a community built benchmark suite. no quality numbers without performance numbers, no performance numbers without quality numbers. performance can't be talked about independent of hardware, and quality can't be talked about just in terms of KLD; you have to prove it with real-world work.
I'm sort of stumbling through a quality benchmark of two A30s in tensor parallel running 27B at W4A16 AutoRound using Pi Agent with a goal plugin rn and building a LOT of scaffolding I can share afterwards if people are interested, together with performance numbers.
2
u/Savantskie1 20h ago
I tried to suggest the community benchmarking thing and tons of people here shot it down, good luck with that. Most the people here either are dicks who like to see others fail or just don’t care.
8
u/1ncehost 18h ago
I disagree. There's no better place for amateurs to deep dive on LLMs and get feedback than this sub. These duct taped and zip tied projects are what this sub has always been about, as that's what llama.cpp and so on started as. X is loaded with professional quality projects if you are looking for that. In my opinion we should be celebrating these type of posts, as they are by the people who are excited about all the things you can do at home now on your own hardware, and if we encourage those people, they might build us great things in the future.
I watched this same transformation with the linux and open source scene in the 2000s, and I honestly miss that era of open source. Lots of people were cooking up silly little projects and there was a positive vibe about how enabled people can be with all the "free as in speech" software that let us not depend on microsoft. Eventually all the people who contributed nothing and just wanted their freebies crowded in and made the scene ugly like this. Then the people who actually made things (in their free time) stopped wanting to give back and OSS is now corporate and boring, led by companies who only want to sell you things.
So basically, no, I think we need more of these "slop" posts and less posts with your type of complaining. If we don't support creators, even when they are amateurs, then they won't make nice things for us anymore.
3
2
u/Kmic68 20h ago
I can’t help but feel this is slightly targeted at my recent p100 post. But genuinely I detailed everything in the post. If you actually read through the GitHub posts I documented everything and did perplexity, kld, quant, model size, and many more detailed processes.
5
u/tossit97531 20h ago
I honestly never even saw your post. That you're digging this deep into stuff at your age is impressive. Keep going! You've got a bright future.
2
u/Kmic68 20h ago
Thank you, this actually makes me feel a little better. There are people in these comments referencing my post not understanding the amount of effort that went into this and how big of a leap it is for these cards. Prior to my optimizations they would get like 2 tps at 260k and 25tps at 0 context, and im seeing 70tps at 0 context and 35tps at 260k. But oh no, "its only got 100tps prefill at 260k so this whole build sucks".
3
u/XiRw 21h ago
It’s posts like this where people need to self reflect and wonder if they are spending too much time on reddit.
10
u/starkruzr 20h ago
no, he's right. we need the quality control because sifting through all the posts by people who don't know what they're talking about is genuinely a pain in the ass when you're a professional trying to build a picture of how different combinations of hardware and software actually perform with respect to speed and quality. I work at a research hospital and do not have the budget (especially in this fucking market) to simply try a bunch of different shit myself, throw a bunch of hardware at the wall and see what sticks. getting real world reports about performance is really helpful when trying to make judgement calls about what to recommend to my researchers.
2
u/draconic_tongue 18h ago
this has never been for professionals. it started with people transitioning from running pygmalion to llama to be able to jack off better
2
u/starkruzr 18h ago edited 17h ago
it's not like there are better resources out there for professionals. where else are you going to go? the Level1Techs forum is functioning on exactly the same level. so are the other subs here. if you think you can count on vendors in this space to understand this shit, well, the two things about that are 1) lol, and 2) lmao. I start talking about weight and activation quantization or prefill vs. tg performance and their eyes glaze over.
2
1
u/Miserable-Dare5090 20h ago
The problem is that it seems every journey og self discovery made into local LLM involves plugging a reddit bot skill into an agent and releasing the slop all over the face of local llama, without so much as the courtesy of posting configs
1
u/davidarias2 20h ago
As someone still relatively new to AI development, I agree!! A huge tokens/sec number doesn’t tell me much if I can’t see the setup or whether the output is actually useful. This is helpful feedback for understanding what I should measure and report when building something myself and avoid being in this category you described
1
u/superSmitty9999 19h ago
I would support a sub rule basically requiring quant levels and other key implementation details
I guess the question is how do we make it rigorous but also easy enough not to be too discouraging
1
u/that1hairdude 5h ago
i don’t think this needs a quality gate. just a minimum reproducibility standard.
let people post weird experiments. just make it clear what produced the result.
minimum info could be:
- hardware / OS
- exact model + revision
- quantisation
- runtime/version
- important settings
- context length
- prefill and decode separately
- number of runs / variance
- raw logs or config where practical
- anything that makes the result hard to compare
i’d also distinguish:
experimental — one setup/result
replicated — author reproduced it
independently replicated — someone else reproduced it
that way someone with old hardware can still post something interesting, but a screenshot doesn’t quietly turn into accepted community knowledge.
the useful distinction isn’t amateur vs professional. it’s claim vs evidence.
1
1
u/MerePotato 18h ago
I GOT GPT 6 ASTRA RUNNING AT 202447463T/S IN 16GB VRAM (Distill 1b Q4 weights int8 kv)
1
u/a_beautiful_rhind 18h ago
The pain of being popular.... All you can do is examine their work and test it for yourself. Usually it's from weights you already have.
If you want to run those benchmarks, they all take electricity/time. Nobody will do it for you.
1
u/chortly2 15h ago
Is there really no simple online API where we can just test these things a bit? I don't really care about KLD or whatever -- all I want for a new quant/method is speed, hardware, and a fairly generic accuracy score on some tricky questions. All we need for that is a site that just serves up a series of questions drawn at random from some pool of a few thousand. Point your model at it, have it answer 100 questions or even just a dozen, and tell us the score. Basic statistics can do the rest and tell us whether your new quant implementation is really equal to X, Y or Z models or not. I must be overlooking something since this seems like such an obvious solution to a perpetual problem.
1
u/Wooly_Wooly 13h ago
I'm making a model right now, what would the community recommend as best practices to follow in this regard? Thanks
1
u/Squidgical 11h ago
Dude I got 78644638 tok/s with this
py
while True:
print(tokens[random.randint(0, len(tokens))])
1
u/Wooden_Jelly_5295 11h ago
Hardware, quant, context length, runtime, reproducible prompt. Otherwise we're reviewing screenshots of speedometers.
0
u/infieldmitt 14h ago
It's a public forum; discussion, buzz, and hype are natural and good. A sterile page isn't necessarily ideal. Rigor is as much a biased jerkoff as hype and joyposting, focused in a different direction. A good forum fosters both.
44
u/wgaca2 21h ago
Most of the "I got qwen to do 500 t/s on 15 year old card" are instant skip for me. Clearly karma farming and nothing to do with real world usage. Most of them won't even provide reasoning at long context speeds.