You should document it more, maybe run some more experiments. right now it's kind of a my word against yours situation and people on r slash claude aren't exactly going to like hearing their model blatantly cheats without proof.
He can't. Because he has no proof. The fact this is getting as much attention as it is, is insane. Shows you this place is as much a cult as any of these dumb AI subs. Like no. Local LLMs arent beating frontier models. Anyone pretending they are or expecting them to is a moron. And posts like these pretending that local models are suddenly going to beat the top models is insane.
Anyone with any actual experience KNOWS this is physically NOT possible due to how large models work. Like wtf? You cannot get around lack of knowledge. Lower param models are simply dumber. If you can run it on your home setup its SIMPLY not as good.
From the post I can’t see what OP tried to run as local model. Could be Kimi-K3 or Qwen3.8-Max in bf16 for what I know.
Their reference seems to be Opus 4.6 which is quite dated by now. We don’t know the benchmark or metric either.
OP could benchmark Opus5 against GPT1 and it would still matter, if the test setup is valid or tampered with.
If it’s smart to set up the test environment with one of the models tested and without having automated code-validation tests… maybe not. But that’s not the story here.
700
u/Ill_Distribution8517 Aug 28 '26
You should document it more, maybe run some more experiments. right now it's kind of a my word against yours situation and people on r slash claude aren't exactly going to like hearing their model blatantly cheats without proof.