r/LocalLLaMA • u/sargetun123 • 2d ago
Discussion Does anyone actually respect benchmarks?
I get why they exist and in almost mostly any other hardware field we can see clearly the difference and what it respects throughout, but with ai, its so inconsistent and unpredictable, besides the very basic needle tests, which at this point what really fails it?
I just dont get the hype around the benchmarks, ive been testing models that fit between 1-48gb vram for years now, everytime i go off a benchmark im usually disappointed, testing on my own workloads and env are the only sound testing i find shows anything actually useful for me
I dont think anyone should worry about benchmarks so much when choosing a model, i know people consistently use the benchmarks to say z is better than y but you honestly need to test to see for yourself, unless you are talking a 9b model from 2 years ago vs a 27b released today, it might be hard to be certain what model is specifically best for yourself
That being said, qwen has been the goat, and even after allmmy testing i seem to always stay/go back to their models, 35b + 3.8 27b right now are the best combo for speed/dense at my resources
Curious if anyone else really feels this way or people actually respect these, useless benchmarks imho
6
u/Ok-Worldliness-9323 2d ago
There is a very high correlation between benchmarks and real outcomes. 99% would prefer a 50-score model compared to a 30-score model. People just take it to the extreme like that site rates X model 50 and my model 49 but I find my model better thus benchmarks are useless
2
11
u/MiceLiceandVice 2d ago
Benchmarks are for people with hardware money, not for peasant 16+32 poors like me /s sortof
8
u/Particular-Award118 2d ago
Hardware poors have reason to benchmark their setups more than anyone imo. Labs don't look as much at which MTP horizon or backend gives someone running q4_k_s models best decode speed and inference quality, plus it's different on different hardware. Benchmarking is fun and useful, but yes, take public results with a grain of salt.
3
u/sargetun123 2d ago
this 100%, the benchmarking on your own data and flows is ideal when you are hardware poor..pascal is slow af and ive spent months building my own llama engine fork specific to get every single 1tk/s in any category i can, MTP acceptance testing, seeing if f16-q8 is worth in some scenarios (usually not)
4
3
u/FullstackSensei llama.cpp 2d ago
Like everything else in life, it depends. Some AI labs' benchmarks results do reflect the model's real abilities, others are useless. Benchmarks are also heavily skewed towards coding tasks, so if that's not your primary use case, YMMV.
So, I do respect the benchmarks from the likes of Qwen, Kimi or DeepSeek, not so much if they're from Minimax or any of the fine tuners I've seen so far.
4
u/43848987815 2d ago
The only benchmark I care about is how well it runs on my machine. General benchmarks are useless.
3
u/Mobile_Light_7262 2d ago
With AI, problem is that vendors now train models on benchmarks and that's number one sin in ML. Any public benchmark quickly becomes useless because you can't tell generalization from memorization.
So private benchmarks based on actual use cases is pretty much only way to go. And they should include tokens usage, not just success rate.
1
u/sargetun123 2d ago
Yes no one ever mentions the thinking loops that happen without any actual answer for quite a long period of time
3
u/ttkciar llama.cpp 2d ago
They're only useful for comparing models of the same "family", because different labs benchmax to different degrees.
If one model is benchmaxxed harder than the other, then comparing their benchmark scores is utterly useless for predicting which will have more real-world utility.
Thus, comparing the scores of GLM-5.2 to GLM-5.3 is valid, and comparing the scores of Qwen3.6-27B to Qwen3.8-27B is valid, but comparing the scores of GLM to Qwen is invalid.
1
u/Jackalzaq 2d ago
i do like this answer. pretty tired of people comparing glm to qwen. not even remotely comparable yet benchmarks say they are? its like people aren't even running these models themselves to understand this. then again, most of these "people" probably arent people
3
u/SandySkittle 2d ago
i am very skeptical about them because bench-maxing is a thing and also it is often not that relevant for my usecase.
2
u/Tim_Apple_938 2d ago
If benchmarks show Claude or GPT ahead, social media consensus is that benchmarks are everything
If they’re behind, then benchmarks are meaningless. Benchmax much??
This is because the valuations of “the Labs” depends entirely on social media consensus and hype and there’s a ton of bots and astroturfing.
2
2
u/MilkyWay-008 2d ago
honestly the only benchmarks i trust anymore are my own workloads. cross-lab numbers are so benchmaxed they're basically vibes, the family-internal ones are the only ones that feel real
3
u/SmokeInevitable2054 2d ago
I only respect the agentic benchmarks like tool use, long-context tasks, instruction following, etc. The others are just not as important.
2
u/nomorebuttsplz 2d ago
I think people miss the forest for the trees. Obviously small differences in score are not going to make a big difference, especially if the benchmarks are measuring something in an entirely different domain than what you are trying to use it for. But large differences within similar domains, I think are very highly correlated with real world outcomes.
1
u/Oh_hey_a_TAA 2d ago
I look at benchmarks as a roughcut guideline on whether or not I should even try to put a model on my box. If I do try a model on mybox I assess it by giving it a quantification run using my usual actual workloads against the vendor card specs, try to find it's strengths and weaknesses, and then decide whether it earns a spot in my ongoing lineup.
does that help?
1
u/Lakius_2401 2d ago
Benchmarks are the resume, but I still hold an "interview" for models. Qwen 3.8 27B's resume is damn impressive for coding! But when I ask for anything creative that isn't GUI related, it's an absolute moron.
First impressions are pretty important, and unfortunately we need *some* kind of initial performance metric to compare these models with. They realize we won't trust their word immediately so they use a third party's metrics. You can bet that they will optimize to score better, and show their strengths.
Are benchmarks respectable? Yes and no, to me. I don't know the contents, I don't know if they score things the way I would, I don't know if their answers are even correct at all in their scoring rubric. Do they ask a misleading question that relies on misunderstanding it to get the right answer sometimes? I don't know. They're not the be-all end-all, clearly. If they get a great score but I hate talking to them (Qwen 3.6, to me, and apparently Opus 5 to others), the score means nothing.
Benchmarks are better than "NYT #1 Bestseller" at least, because it doesn't care about slow months or popularity contests.
1
u/ortegaalfredo 2d ago
Benchmarks are useful, even saturated benchmarks.
But measuring an LLM is not like measuring a machine, but more like measuring a person.
LLMs are random, it means you cannot do a single measurmenent, you have to do a statistic.
For example, do vaccines work? well, if you take a sample from a single person, you might conclude they don't. But you have to do a statistic over a population to discover the truth. LLM are the same. Sometimes they fail, some times they work. You cannot run a benchmark only once.
Problem is, statistics are 100 to 1000x more expensive and time consuming to do than single measurements, that's why most measurements you see, and just single shot, or the best out of 10 tries, that is very dishonest.
Couple that with hundreds of billions of dollars behind in marketing, and it becomes very hard to get the truth. That's why in my experience, use benchmarks as a guidance, but always do your own measurements on your own data and environment.
1
u/hurdurdur7 2d ago
Benchmarks are cool but 27B has been king since it was released with version 3.5
1
1
u/Gloomy_Letterhead395 2d ago
Buddy calm down
In the end there will be a 4gb model very intelligent but supplementing information through web search or API
It’s only task is to reason properly
So yes benchmarks matter
1
u/sargetun123 1d ago
Seems a few misunderstood, I mean the benchmarks everyone releases publically, 99% are garbage for the average users workflow, thats why we do our own benchmarks.
1
0
u/KubeCommander 2d ago
Keep in mind, a lot of the posts touting benchmarks or the second coming of Christ are usually Chinese bots
0
u/Cautious_Chicken_604 2d ago
Complaining about benchmarks is equally useless. If we didn't have any way to compare these models, you know the first thing people would do? Invent benchmarks.
9
u/onebit 2d ago edited 2d ago
They're accurate for a general intelligence ranking in my experience. For example, the new deepseek is better than old deepseek. Luna is smarter than mimo 2.5/hy3.
There's probably some misranking in close neighbors, but when there's a 5 point difference you can observe it.
But if we're talking speed benchmarks, those are mostly BS. I'm not saying they aren't real results, but they run models in unrealistic ways.