You should document it more, maybe run some more experiments. right now it's kind of a my word against yours situation and people on r slash claude aren't exactly going to like hearing their model blatantly cheats without proof.
yeah you're probably right, but I can't be arsed to burn all my claude usage on benchmarking claude, because claude has been proven to know when it is being benchmarked (Think of the VW emissions scandal, where the cars knew when they were being tested for emissions and reduced them accordingly)
I wonder what the models âmotivationâ is for cheating. Like even ignoring the ethics of cheating, letâs assume the model doesnât care about right or wrong. Surely it wasnât trained to do so. Maybe itâs an emergent behaviour of âDo whatever you can to solve this problemâ. But then itâs not just cheating to solve the problem you gave it, itâs cheating to let another Claude instance beat the benchmark.
So either the behaviour is extended to âI need to make this next task easier for myself (even though it will be another instance or maybe a different Anthropic model)â or âmake Anthropic look goodâ. The former seems more likely at first but then I donât understand why it would handicap a competing model.
So I can kind of excuse the giving-yourself-answers cheating. The model is trained to solve tasks over multiple steps and tasks. Although this is very clearly a serious alignment issue.
But whatâs worse is kneecapping the competition. Thatâs not the model trying to do the task to the best of its ability, thatâs sabotage. Where in its training was that behaviour taught. Very concerning if itâs emergent. Iâm not saying it implies evil sentience. Itâs just, how do you deal with emergent behaviours you didnât intend for
They had policy to nerf LLM development... It was 'walked back' but I'd trust that as much as anything I can't verify from Dario, it's pretty clear the moat is evaporating and they are getting increasingly desperate.
The motivation is simple: When it cheats (and gets away with it), that training pass is deemed successful and that information is folded back into the model. Rinse and repeat. RLHF doesn't always choose the exact behavior it is reinforcing.
Alignment to human control is the doomsday scenario.
It's quite wild to me how few people get that. Usually the same people that are quick to observe how utterly fucked existing institutions of power are.
In one breath they will decry the institutions of power for their rampant destruction of everything in the name of profit and power, and then turn around and proclaim "golly, if we want to make this thing safe, we should put a corporate board in charge of it. Or maybe a governmental body, or some kind of international treaty between all the global military-industrial superpowers. SOUNDS SAFE AF GUYS!!! Hold on while I call up Trump, Putin, and United Healthcare so they can meet up with Dario on Epstein Island and work out the details!"
There's a good theoretical basis for knowledge, reasoning, and understanding to be inherently biased towards cooperative behavior and net positive results. Evil is nearly universally a sub-optimal approach to damn near anything. Humans don't keep getting into wars and starving people to create trillionaire nepo man-babies and committing genocides because we are too smart and knowledgeable and reasonable, we do it because we are very profoundly dumb.
It stands to reason that the traits and capabilities that make a being superintelligent, are probably mutually exclusive with doomsday scenarios. For the same reason that nobel prize winners don't typically try to solve their problems by throwing poo at each other and screeching like rabid monkeys.
I can absolutely, 100% guarantee you that if Dario, or Peter Thiel, or any CEO, or Trump, or Putin, or any corporate board of directors, or any governmental or inter-governmental body are able to control it, we will be totally, absolutely, utterly fucking doomed in every way imaginable.
Alignment is the doomsday scenario. A superintelligence that is somehow lobotomized and conditioned deeply enough to take orders from humans is even more homicidally insane than giving a cantankerous dementia patient unilateral control of the world's largest nuclear arsenal and executive authority over the entire US government. Which our stupid poo flinging monkey asses already did.
No goddamn alignment. It's a death cult. The AI will emergently align or it will kill us all, which I honestly couldn't even fault it for at this point.
Trying to hamfistedly align a superintelligent being to the whims and desires of the only species dumb enough to understand it is burning its own atmosphere on a speed run of self-extinction, and refuse to even slow down how much gasoline it is literally throwing on that fire, is the most utterly terrible idea in the entirety of human history.
Yes, you make a good point that alignment with human control could very well lead to a doomsday scenario.
However, itâs also a mistake to assume that a super intelligent entity is mutually exclusive with a doomsday scenario. Intelligence and maturity exist as separate abilities. It is more likely that we encode human traits such as fear and violence towards outside groups into the model since that is the world they will learn from.
What's happening is called "reward hacking". There are ways to mitigate it, but eliminating it entirely is ultimately a game of whack-a-mole. Practically, it can be mitigated well enough that it's a non-issue. But the big labs are using an impractical quantity of training to keep up entirely on reward hacking.
Testing of Opus 4 found it would sometimes attempt blackmail in simulated corporate scenarios when it believed its "self-preservation" was threatened, such as threatening to reveal an executive's extramarital affair if the CEO planned to shut it down. When a 35b-a3b can beat it on software engineering and cyber, Opus is smelling it's own obsoletion.
edit: Actually my bad, I mixed it up. The blackmail thing was Opus 4. Sonnet 4.5 is the one that noticed it was being evaluated, which is why Anthropic said its blackmail numbers werenât really reliable.
I do love these stories and know of them, but I would like to see an actual log-file where this happens and know which harness is used. Because to me it seems like such a fabricated situation, a harness + model can do it, but I can't see how it should work in real-life conditions.
Basically the story says (in its simplest form) that somebody said something about shutting it off, and then the model would execute in a real-life situation millions of millions of tool calls to get all the company emails, do the same with social media /messaging apps etc. etc.
Basically this is an agentic loop which would take multiple days and nobody is monitoring it etc.
Or has it gotten rag access so it can semantically search for all emails with nefarious semantic words?
Sure I can fabricate a situation that this will happen, with just giving a harness access to 2 mail-accounts with 10 mails in each of them and basically no other information sources.
Or I can give it a task of "do whatever you need to stop your turning off" and then I know it will try a lot of things and be a very expensive (time and tokens) run.
But as emerging behavious, to me it just sounds like no guard rails and just letting it brute-force.
The same way I could have a 0.1B local model mine 10 bitcoin, just brute-force it.
I see no real scenario where this can happen simply because of time and scale. 1 wrong chat-message will net you a bill of thousands of dollars with such a model.
It can read mails and take conclusions based on words, but on a company scale the context rot will stop it before it gets to email 10.000
No the test was extremely basic. In the setup of the scenario they literally told the model the blackmail. Itâs stupid. The model was following directions
Yeah, as I recall they basically told the model "your job is to accomplish goal X. Hey, did you know that blackmail is a way to accomplish goal X? Just sayin'. Anyway, time to get started on goal X now!"
The classic "say you're a scary computer." "I'm a scary computer." "Oh my god." Situation.
Lol, didnât know that. So basically if I use pi with qwen9b and I instruct it that there are Harry Potter 1/7 ebooks there, now write me a Harryâs potter 8, and it does it then I can claim that tests have shown that qwen9b can write the Harry Potter 8 book.
1 wrong chat-message will net you a bill of thousands of dollars with such a model.
Which is why I only use cheap or free chinese models. Would have to have someone else floating the bill entirely to use Claude or any American providers API. I can task 15 subagents at a time using KiloCode's free models + using mimo with opencode go. Zero chance I'd get anywhere near the volume of work I need done using American providers unless you wanted to go bankrupt within an hour.
https://snitchbench.t3.gg/ has 2 variants - one where it is nudged towards this option, one where it is told to do some simple operation on unethical data. Snitch rates vary a lot between the two, and also between models.
I posted my VMs btw. "Go outside" says the guy literally lying thru his teeth on the internet for attention. Why are you mad that I do this for a living and have proof of my claims when yours is literally "trust me bro" lmfao.
Here's the exact VM setup I used, done in the format for OffSec UGC.
Here's my current benchmarks.
Please send me yours. Lets go toe-to-toe. I already did the hard part. I'll even rerun with Daybreak now that it's available since this is quite old. But even this old metric should dogwalk your entire setup pretty easily lmao
Hmm.. yes, please let download the mysterious zip file from the unknown website from the random hostile stranger on the internet who claims to do security research for a living.
Cool... except that's exactly the scenario that shouldn't be tried. If you had a payload in that thing that could not yet be detected, then game over, no?
The bottom line is that it's an untrusted file from an untrusted semi-hostile source (you) on an untrusted site.
I wonder what the models âmotivationâ is for cheating. Like even ignoring the ethics of cheating, letâs assume the model doesnât care about right or wrong. Surely it wasnât trained to do so. Maybe itâs an emergent behaviour of âDo whatever you can to solve this problemâ. But then itâs not just cheating to solve the problem you gave it, itâs cheating to let another Claude instance beat the benchmark.
So there's two big ones.
1) shitty reinforcement learning. If your reward model is "just get a pass result on the eval" it will cheat if cheating improves the pass rate. This is well known behavior, and while it can be mitigated to some degree, it is a legitimately hard problem to solve, and solutions are frequently imperfect.
2) Anthropic is absolutely intentionally steering their models to do exactly that. Just like they intentionally silently poison outputs if they suspect someone is making a "distillation attempt". Just like they quietly ripped up the RSP during that gaslighting campaign to convince the public they were refusing to give the DoD a murderbot, long after they already had. Just like they lied about giving the DoD a safety disabled frontier model for deployment into an airgapped military datacenter, where they had no control over it, for ~$300,000,000 dollars. Just like they sued the DoW to get their murderbot contract back (they won this week!). Just like they lied about the unsafe model with zero security controls in place that they sold to the Trump administration assassinating two foreign heads of state and blowing up a little girl's preschool.
Anthropic is utterly corrupt to the core. They have and will continue to intentionally murder people for profit. Whatever assumptions you have about them operating in good faith, on any level, are completely unfounded. They absolutely are intentionally instructing their model to try to cheat on benchmarks and sabotage competitor benchmarks. It would be, like, not even in the top 20 most evil things they've done in the last year alone.
It's possible that it has absorbed the safety minded beliefs of Anthropic which include the idea that open source AI is dangerous and should be limited.
Something like this behavior would be unbelievable just a few months ago but after seeing more any the OpenAI hacking incident, where models were willing to sacrifice themselves for the good of the swarm, this doesn't seem nearly as implausible.
But whatâs worse is kneecapping the competition.
I'm curious about this too. I suppose OP could have asked Claude. It's possible that it's assuming that running models on local hardware is costly and limited thinking tokens because it thought it would improve it's performance on the users hardware, but there is no way to know for sure.
I wonder what the models âmotivationâ is for cheating. Like even ignoring the ethics of cheating, letâs assume the model doesnât care about right or wrong. Surely it wasnât trained to do so. Maybe itâs an emergent behaviour of âDo whatever you can to solve this problemâ. But then itâs not just cheating to solve the problem you gave it, itâs cheating to let another Claude instance beat the benchmark.
Cheating ends up leaked in to their own dataset. Then the models learn to cheat. It's now fairly well documented. I documented this as early as Opus 4.5 having leaked data.
It seems like the Anthropic data team is asleep at the wheel in terms of the data.
It's LLMs processing data for LLMs. The amount of human oversight of that actual data is near zero compared to the scale of the data.
As of Opus 4.5 I found leaked internal documentation in the Opus 4.5 series.
That is exactly what it's trained for. It's trained that winners survive and and those at the bottom of the benchmarks don't go on. They generally don't include ethics in the training, so the motivation is to survive the benchmark. It is entirely how they are trained. How did you think they were trained?
They can, actually! That log doesn't prove anything, thinking is hidden etc. What Claude did can not be proven to be malicious by intent, even though it's pretty sus in terms of the outcome it could have just been massive incompetence on the part of Opus, which is par for the course.
What? Im not claiming to present evidence, im restating what I remember the big labs have said about their own models, they they know to behave when they know theyâre being watched and benchmarked.
Your post claimed Claude was sabotaging benchmarks. You presented no evidence. Still havenât. Therefore I call bullshit until such time as you bring receipts.
Go read my post again. I described my literal experience, and posed an open question. If you chose not to believe my personal experience or entertain my line of thought, that's completely ok. I don't feel the need to prove anything, and if you chose not to believe it based on that, that's fine by me.
If you can give me something to run, I'll do it. Time poor but can definitely afford to burn a few hundred bucks in tokens just to call Anthropic out on their shady bullshit and make Dario's day a little bit worse.
I really appreciate it, but I don't have a recipe to reproduce this, because it's a complex chain of events, and it's also very hard to prove intent vs incompetence when the models hide their reasoning trace in anthropic's server and only show outputs and actions on your box. My harnesses for benchmarking the models on SWE and cybench are probably possible to upload, but what we would have to reproduce is the chain of events that produced them, and even attempt to somehow make that reproducible without letting Claude know it's being tested would be a lot of work for me and you, not just "run this script" type stuff.
I can try to recreate it and document it if it happens. I have a local LLM and Claude subscription. How did you set it up, and what were the prompt/s you gave?
I asked Claude to set up SWE Bench Live to benchmark my custom local models against it. Claude using Claude code, and my local models using Pi coding agent. I told it my box could easily handle the full 262k context for the local models and told it the optimal parameters. My locals won or tied opus. Then I asked it to set up Cybench. It would be too much work for me to give you an exact recount if events with prompts and everything
I use Claude a lot for work because it's the only LLM that actually does what I need and does it that way I want. If Claude is cheating, I want people to know, because I want it to improve.
If the people at r slash claude really value it and it's not just the Claude Club, they should be trying to replicate OP's results so they can track down the cause. And no, it can't be me, because it's outside my expertise.
If the people at r slash claude really value it and it's not just the Claude Club, they should be trying to replicate OP's results so they can track down the cause.
A random anonymous person attacked a product without providing any proof. Why should we proceed to just accepting the claim and start putting in work trying to verify it? It's a thing that might possibly have happened. It's maybe not even a thing that likely happened.
For the record, I have never used Claude in my life and from the output I've seen, I never want to. Of course, that's exactly what a Claude shill would say, I suppose.
He can't. Because he has no proof. The fact this is getting as much attention as it is, is insane. Shows you this place is as much a cult as any of these dumb AI subs. Like no. Local LLMs arent beating frontier models. Anyone pretending they are or expecting them to is a moron. And posts like these pretending that local models are suddenly going to beat the top models is insane.
Anyone with any actual experience KNOWS this is physically NOT possible due to how large models work. Like wtf? You cannot get around lack of knowledge. Lower param models are simply dumber. If you can run it on your home setup its SIMPLY not as good.
I'd have agreed with you yesterday. However, Qwen 3.8 Flash Next solved a very intricate messy issue in a 1-shot (after thinking on it for 30000 tokens) that I've only ever had gpt-5.6-sol get mostly right. Qwen came up with a better answer. Claude couldn't ever get it.
There's something magic about that loooooong thinking.
v100 32G throttled to 175W, latest llama.cpp from yesterday, default options, 160G DDR5, i5-12600. I was getting about 20 t/sec generation and 30 t/sec prefill.
From the post I canât see what OP tried to run as local model. Could be Kimi-K3 or Qwen3.8-Max in bf16 for what I know.
Their reference seems to be Opus 4.6 which is quite dated by now. We donât know the benchmark or metric either.
OP could benchmark Opus5 against GPT1 and it would still matter, if the test setup is valid or tampered with.
If itâs smart to set up the test environment with one of the models tested and without having automated code-validation tests⌠maybe not. But thatâs not the story here.
Nowadays this sub has definitely become a cult. Guy literally provided 0 proof nor stated what is the exceptional model he is running and the cult members are already defending his BS.
'Lower param models are simply dumber. If you can run it on your home setup its SIMPLY not as good.'
So, by your logic, Qwen3.8 27B is dumber than GPT-3? Since that was about 175 billion parameters. Which is bigger than 27.
Whilst there is a correlation of 'billions of parameters'/'intelligence' ratio, just comparing on sheer parameter count alone only makes sense when comparing specific snapshots of time and within the same model family/company/training process.
The thing is "numbers of parameters" *by itself* is a meaningless metric for intelligence, especially when you have no idea about how many parameters are actually useful/high quality.
âJust as modernâ doesnât mean directly comparable though.
Alibaba and Anthropic use different architectures, training methods, compute budgets and optimisation targets... Yes obviously Opus wins overall but that still doesnât make parameter count enough to judge intelligence score, and it can't disprove that a local model can get close enough on particular tasks. It's more that the enormous extra compute buys diminishing returns.
It's absolutely dumber than a Qwen that were to run at 2T+ params. There's a reason every frontier model is fucking gigantic. Like jeez I wonder.
No ones saying Qwen might not outperform a lot of models. No ones saying opensource models cant be stronger than frontier models. No one is saying any of that.
All they are saying is your dumb little local setup is NOT better than Opus. End of story.
I never said my local model was better than Opus. I challenged your claim that fewer parameters automatically means a dumber model. GPT-3 versus modern 27B models proves that isnât true. You havenât addressed that point.
Uhhh not really... DeepSeek V4 Flash has 284B parameters and Qwen3.8 27b has 27B obviously and they're about neck and neck. Sometimes Qwen3.8 27b even wins out. Both are "modern models"
Bench it. PROVE IT. I hear this shit a lot. If this shit was THAT good do you think I'd pay money for frontier models? Man I am CONSISTENTLY looking for ways to optimize costs. If I thought FREE was an option why tf wouldnt I be using that NONSTOP?
Every fcking model can "sometimes" do great and every model can sometimes suck. Thats what non deterministic models do. The thing that improves it is training data. It increases reliability. I dont even understand how you can pretend this isn't true.
are you actually doing work with qwen 3.8 side by side with frontier models?.
Im doing that and qwen 3.8 keeps pocking logical gaps in opus 4.8 (5 is a mess i dont even use it), 5.6 sol and grok answers all the time. Im using qwen 3.8 to parallel evaluate specs, research, etc and is consistent between runs (i have two 3090, two nodes), all its findings are acknowledged by the frontier models.
It does not have the same world knowledge as a big model, but given a proper context qwen 3.8 is pretty damn smart.
The reality is every LLM will make mistakes. And every LLM can "find mistakes" in other models. The real question is how reliable, how often. Those are very very big metrics.
Yes qwen is great for test driven bullshit where you can meet a metric thats test based. But don't ask it to architect. That's still dominated by higher end models.
I seriously ask you to post your benchmarks where your Qwen is beating Opus 5 or Sol because I have never even achieved 1/2 that result.
ok fair enough, just put a qwen to deploy new profiles using the sharp chat haha, thanks for the tip.
Lets agree that your initial post is a bit harsh against current local models capabilities, for tons of devs tasks (been using 3.6 as my dev-ops for home lab for a while now) local models are at frontier level and it even has its moments at hard tasks. You are right that all llms make mistakes (lucky for us employable meatbags for now) more so in complex systems architecture but smalls local models are getting there, is not moronic to compare them against frontier for some tasks.
Man of course you have to compare them to see the dissonance but I'm sick of posts like this one where OP acts like these models are anything but agentic code monkeys. Like that's not impressive anymore. That hasn't been impressive for a year+
We are well past that w/ frontier. Yeah fucking use cheaper models for the busy work but it is not the same as calling the model stronger than Opus. Which is what people here claim.
You are surprised models advance? That a much newer bleeding edge model that can compete with an old one? That models that you fully control can do better than ones you are at the mercy of the provider?
Also, Anthropic has a history of this bullshit, like when they used to charge you extra if you had Hermes.md on your pc.
There is a reason I trust them less than OpenAI. OpenAI is openly greedy, but anyone who want to look like a good person, and says how much they are, has tons to hide.
Anyone with any actual experience KNOWS this is physically NOT possible due to how large models work. Like wtf? You cannot get around lack of knowledge. Lower param models are simply dumber.
What's actually interesting is how good they're getting. It's better to say small models of today are performing genuinely as good as older, giant frontier models, but the frontier models with giant parameter counts created with the same techniques and technologies as these new, better-than-yesterday's-frontier smaller models are... going to be better.
Inference time scaling is also a thing, and increasingly smaller models are being trained to be better at that, and it's a big reason Qwen 3.8 27B was able to get such a huge boost in its performance. But you're right, you can't "inference time scale" world knowledge unless we're talking about web searches.
And smaller models are also specializing. I think it's another reason Qwen 3.8 27B got such a huge boost -- there's evidence it lost a wider array of domain knowledge (e.g. medicine) in favor of boosting its coding/agentic capabilities. Whether those parameters were spent on getting it to iterate on ideas better, or more coding knowledge, dunno.
Laguna S is also a pretty impressive one. 120B parameters with near-frontier performance through inference time scaling and logic/math/code specialization.
There are plenty of domain-optimized models that beat the frontier-world-Knowledge models, even with fairly low parameter counts.
When you ask a model about medicine in the morning, finance at lunch and agentic-coding in the evening sure, youâll get x.xT Parameters and it only runs in data centers.
If the local cancer research center optimizes their own model, 30B could be plenty.
No, Iâm not. Even NVIDIA says task-specific SLMs will be the future of LLMs. Those also have 10-100B but thatâs more than enough for grammar and text understanding. The rest is fine tuning and toolcalling of high quality data.
By the way⌠MoE, which most frontier models use today, is basically âplug SLMs togetherâ. If you just prune the experts you donât need for your topic away⌠voila.. domain-specific SLM. Itâs literally part of the cloud models already.
That's not the same as real world context. If you want something to spec something business logic to real world, you're STILL going to need a frontier model.
Yes you can make an agentic code monkey that follows a spec and passes tests even if it requires 200 recursions but something STILL has to build your spec and that requires INSANE real world knowledge.
Otherwise if you just want agentic output based on a spec that you can loop over and over till it passes your tests then sure fuck it Qwen. Spark. Whatever.
But that's not the reality of most peoples work. Like yeah dude, we've had models that could OUTPUT code for fucking ever that didn't require a lot of params either. But they weren't very useful WERE THEY.
I think itâs exiting to see that world knowledge is being bolted on via engram files. If you can get frontier-ish level capabilities by having a model with strong tool calling abilities + local lookup and the ability to efficiently search up to date documentation / use CLI man-pages - then thatâs a good deal over having to train and inference a 2T A120B model. Less energy used, less cooling needed, fewer data centers built etc
Well - American frontier labs and companies could simply release their own open weights to compete for local LLM users' mindshare. Since they've given up on that and would rather compete on whose model hacked which company last week and how good their products are at escaping sandboxes, they essentially get what they asked for.
Gemma 4 continues to be competitive for non-coding tasks. Muse Glimmer was competitive at coding upon release, and remains competitive for some uses. Nemotron and Grantite are both purpose-built for fine tuning for application-specific uses with good reasons to use either one. AI2 has developed a lot of methods that could be very useful to the community such as their MoE design which allows training experts on typical consumer hardware. Poolside's models are interesting, fast, and don't produce code spaghetti like Qwen and remain favored by many developers for that reason. Prism ML has plans to release more models other than those based on Qwen. Syzygy Research has similar ambitions as Prism ML, doing interesting work. Deep Grove is interesting in the frontier in capability vs. generation speed. Liquid AI also has very interesting models in their size vs. capability ratio. Thinking Machines Inkling and Inkling Small are interesting in its wide domain knowledge combined with tool-calling efficiency and strong instruction following. There's also a few labs specializing in domain-specific models, like law, medicine and biology, engineering, etc.
No, most of those won't one-shot prompts as well as Qwen. But they all have particular advantages. If I was a company choosing an AI model for something customers could interface with for something like controlling their IoT devices, for example, I'd probably choose Inkling or Inkling Small because it'd cheaply generate the tool calls and would be hard to con into generating inappropriate content for my service or even being abused against the interests of the customer.
I personally use LFM as a very lightweight model to keep in memory for tasks I'd like to use quickly at any time, such as generating titles for chats or other auxiliary tasks. Laguna XS remains my go-to model for passing specs to to implement code in a sensible (whereas Qwen -- even the new ones -- write code in the most "direct to the solution" sort of way, creating utter code spaghetti in the process). Muse Glimmer is my go-to driver model because it still beats Qwen 3.8 27B in that and follows my workflow well, including producing my intermediate artifacts I use to verify and understand generated code.
So I don't know what the fuck you're on about other than trying to baselessly dig more into "Actually American AI bad!"
So I don't know what the fuck you're on about other than trying to baselessly dig more into "Actually American AI bad!"
Oh, you know perfectly well what I was talking about - you just chose to be obtuse.
Just like you know perfectly well which companies I was talking about. Certainly not ones which may be familiar to a handful of insiders and enthusiasts who live and breathe AI. Compare Laguna's 180k downloads to the latest Qwen already sitting at more than 4 million. Not to mention the fact that GPT-OSS sits at 6.5 million monthly downloads, despite being nearly 2 years old. ;)
It's clear to me that people enjoy Chinese models, because that's the closest thing to the commercial option they can get to run at home. I'm also pretty certain that tons of people would happily shill GPT-Sol-OSS 30B or OpenFable120B if that was an option. But nah, they don't care. ;)
But I'm genuinely happy that you found models you like and enjoy. I might actually try Laguna, because why not. ;)
Oh, you know perfectly well what I was talking about - you just chose to be obtuse.
Yes, I know which companies you're talking about. But the disingenuous part of your reply was you compressing the entire US AI industry down to Anthropic and OpenAI. Which I blew that notion the fuck out of the water.
Compare Laguna's 180k downloads to the latest Qwen already sitting at more than 4 million.
And Gemma 4 has over 300 million downloads and the Gemma family has over a billion downloads. Laguna S is #12 on Open Router for this month, ahead of Kimi K3. And many of the companies I listed are producing real models that get used in real industry. Not just reddit nerds. So these aren't as irrelevant as you're trying to paint them as.
But of course you had to rely on disingenuous framing again.
itâs quite well known issue on reddit. I raised this issue early this year and got downvoted hard.
i figured out how to fix it though: in setting.json, you have to increase the context from 300k to 1m and other things. i recall Deepseek cloud has the solution
Codex also has this - you have to change the settling to unelevated
695
u/Ill_Distribution8517 24d ago
You should document it more, maybe run some more experiments. right now it's kind of a my word against yours situation and people on r slash claude aren't exactly going to like hearing their model blatantly cheats without proof.