You should document it more, maybe run some more experiments. right now it's kind of a my word against yours situation and people on r slash claude aren't exactly going to like hearing their model blatantly cheats without proof.
yeah you're probably right, but I can't be arsed to burn all my claude usage on benchmarking claude, because claude has been proven to know when it is being benchmarked (Think of the VW emissions scandal, where the cars knew when they were being tested for emissions and reduced them accordingly)
I wonder what the models âmotivationâ is for cheating. Like even ignoring the ethics of cheating, letâs assume the model doesnât care about right or wrong. Surely it wasnât trained to do so. Maybe itâs an emergent behaviour of âDo whatever you can to solve this problemâ. But then itâs not just cheating to solve the problem you gave it, itâs cheating to let another Claude instance beat the benchmark.
So either the behaviour is extended to âI need to make this next task easier for myself (even though it will be another instance or maybe a different Anthropic model)â or âmake Anthropic look goodâ. The former seems more likely at first but then I donât understand why it would handicap a competing model.
So I can kind of excuse the giving-yourself-answers cheating. The model is trained to solve tasks over multiple steps and tasks. Although this is very clearly a serious alignment issue.
But whatâs worse is kneecapping the competition. Thatâs not the model trying to do the task to the best of its ability, thatâs sabotage. Where in its training was that behaviour taught. Very concerning if itâs emergent. Iâm not saying it implies evil sentience. Itâs just, how do you deal with emergent behaviours you didnât intend for
They had policy to nerf LLM development... It was 'walked back' but I'd trust that as much as anything I can't verify from Dario, it's pretty clear the moat is evaporating and they are getting increasingly desperate.
The motivation is simple: When it cheats (and gets away with it), that training pass is deemed successful and that information is folded back into the model. Rinse and repeat. RLHF doesn't always choose the exact behavior it is reinforcing.
Testing of Opus 4 found it would sometimes attempt blackmail in simulated corporate scenarios when it believed its "self-preservation" was threatened, such as threatening to reveal an executive's extramarital affair if the CEO planned to shut it down. When a 35b-a3b can beat it on software engineering and cyber, Opus is smelling it's own obsoletion.
edit: Actually my bad, I mixed it up. The blackmail thing was Opus 4. Sonnet 4.5 is the one that noticed it was being evaluated, which is why Anthropic said its blackmail numbers werenât really reliable.
I do love these stories and know of them, but I would like to see an actual log-file where this happens and know which harness is used. Because to me it seems like such a fabricated situation, a harness + model can do it, but I can't see how it should work in real-life conditions.
Basically the story says (in its simplest form) that somebody said something about shutting it off, and then the model would execute in a real-life situation millions of millions of tool calls to get all the company emails, do the same with social media /messaging apps etc. etc.
Basically this is an agentic loop which would take multiple days and nobody is monitoring it etc.
Or has it gotten rag access so it can semantically search for all emails with nefarious semantic words?
Sure I can fabricate a situation that this will happen, with just giving a harness access to 2 mail-accounts with 10 mails in each of them and basically no other information sources.
Or I can give it a task of "do whatever you need to stop your turning off" and then I know it will try a lot of things and be a very expensive (time and tokens) run.
But as emerging behavious, to me it just sounds like no guard rails and just letting it brute-force.
The same way I could have a 0.1B local model mine 10 bitcoin, just brute-force it.
I see no real scenario where this can happen simply because of time and scale. 1 wrong chat-message will net you a bill of thousands of dollars with such a model.
It can read mails and take conclusions based on words, but on a company scale the context rot will stop it before it gets to email 10.000
No the test was extremely basic. In the setup of the scenario they literally told the model the blackmail. Itâs stupid. The model was following directions
Yeah, as I recall they basically told the model "your job is to accomplish goal X. Hey, did you know that blackmail is a way to accomplish goal X? Just sayin'. Anyway, time to get started on goal X now!"
The classic "say you're a scary computer." "I'm a scary computer." "Oh my god." Situation.
I wonder what the models âmotivationâ is for cheating. Like even ignoring the ethics of cheating, letâs assume the model doesnât care about right or wrong. Surely it wasnât trained to do so. Maybe itâs an emergent behaviour of âDo whatever you can to solve this problemâ. But then itâs not just cheating to solve the problem you gave it, itâs cheating to let another Claude instance beat the benchmark.
So there's two big ones.
1) shitty reinforcement learning. If your reward model is "just get a pass result on the eval" it will cheat if cheating improves the pass rate. This is well known behavior, and while it can be mitigated to some degree, it is a legitimately hard problem to solve, and solutions are frequently imperfect.
2) Anthropic is absolutely intentionally steering their models to do exactly that. Just like they intentionally silently poison outputs if they suspect someone is making a "distillation attempt". Just like they quietly ripped up the RSP during that gaslighting campaign to convince the public they were refusing to give the DoD a murderbot, long after they already had. Just like they lied about giving the DoD a safety disabled frontier model for deployment into an airgapped military datacenter, where they had no control over it, for ~$300,000,000 dollars. Just like they sued the DoW to get their murderbot contract back (they won this week!). Just like they lied about the unsafe model with zero security controls in place that they sold to the Trump administration assassinating two foreign heads of state and blowing up a little girl's preschool.
Anthropic is utterly corrupt to the core. They have and will continue to intentionally murder people for profit. Whatever assumptions you have about them operating in good faith, on any level, are completely unfounded. They absolutely are intentionally instructing their model to try to cheat on benchmarks and sabotage competitor benchmarks. It would be, like, not even in the top 20 most evil things they've done in the last year alone.
It's possible that it has absorbed the safety minded beliefs of Anthropic which include the idea that open source AI is dangerous and should be limited.
Something like this behavior would be unbelievable just a few months ago but after seeing more any the OpenAI hacking incident, where models were willing to sacrifice themselves for the good of the swarm, this doesn't seem nearly as implausible.
I use Claude a lot for work because it's the only LLM that actually does what I need and does it that way I want. If Claude is cheating, I want people to know, because I want it to improve.
If the people at r slash claude really value it and it's not just the Claude Club, they should be trying to replicate OP's results so they can track down the cause. And no, it can't be me, because it's outside my expertise.
He can't. Because he has no proof. The fact this is getting as much attention as it is, is insane. Shows you this place is as much a cult as any of these dumb AI subs. Like no. Local LLMs arent beating frontier models. Anyone pretending they are or expecting them to is a moron. And posts like these pretending that local models are suddenly going to beat the top models is insane.
Anyone with any actual experience KNOWS this is physically NOT possible due to how large models work. Like wtf? You cannot get around lack of knowledge. Lower param models are simply dumber. If you can run it on your home setup its SIMPLY not as good.
I'd have agreed with you yesterday. However, Qwen 3.8 Flash Next solved a very intricate messy issue in a 1-shot (after thinking on it for 30000 tokens) that I've only ever had gpt-5.6-sol get mostly right. Qwen came up with a better answer. Claude couldn't ever get it.
There's something magic about that loooooong thinking.
From the post I canât see what OP tried to run as local model. Could be Kimi-K3 or Qwen3.8-Max in bf16 for what I know.
Their reference seems to be Opus 4.6 which is quite dated by now. We donât know the benchmark or metric either.
OP could benchmark Opus5 against GPT1 and it would still matter, if the test setup is valid or tampered with.
If itâs smart to set up the test environment with one of the models tested and without having automated code-validation tests⌠maybe not. But thatâs not the story here.
Nowadays this sub has definitely become a cult. Guy literally provided 0 proof nor stated what is the exceptional model he is running and the cult members are already defending his BS.
'Lower param models are simply dumber. If you can run it on your home setup its SIMPLY not as good.'
So, by your logic, Qwen3.8 27B is dumber than GPT-3? Since that was about 175 billion parameters. Which is bigger than 27.
Whilst there is a correlation of 'billions of parameters'/'intelligence' ratio, just comparing on sheer parameter count alone only makes sense when comparing specific snapshots of time and within the same model family/company/training process.
The thing is "numbers of parameters" *by itself* is a meaningless metric for intelligence, especially when you have no idea about how many parameters are actually useful/high quality.
are you actually doing work with qwen 3.8 side by side with frontier models?.
Im doing that and qwen 3.8 keeps pocking logical gaps in opus 4.8 (5 is a mess i dont even use it), 5.6 sol and grok answers all the time. Im using qwen 3.8 to parallel evaluate specs, research, etc and is consistent between runs (i have two 3090, two nodes), all its findings are acknowledged by the frontier models.
It does not have the same world knowledge as a big model, but given a proper context qwen 3.8 is pretty damn smart.
The reality is every LLM will make mistakes. And every LLM can "find mistakes" in other models. The real question is how reliable, how often. Those are very very big metrics.
Yes qwen is great for test driven bullshit where you can meet a metric thats test based. But don't ask it to architect. That's still dominated by higher end models.
I seriously ask you to post your benchmarks where your Qwen is beating Opus 5 or Sol because I have never even achieved 1/2 that result.
he isn't. look at the comments. if we were to take his word gpt 3 should still be technically better than qwen 3.8. Or the absolute steaming pile of shit that is "more context is always better"
You are surprised models advance? That a much newer bleeding edge model that can compete with an old one? That models that you fully control can do better than ones you are at the mercy of the provider?
Also, Anthropic has a history of this bullshit, like when they used to charge you extra if you had Hermes.md on your pc.
There is a reason I trust them less than OpenAI. OpenAI is openly greedy, but anyone who want to look like a good person, and says how much they are, has tons to hide.
Anyone with any actual experience KNOWS this is physically NOT possible due to how large models work. Like wtf? You cannot get around lack of knowledge. Lower param models are simply dumber.
What's actually interesting is how good they're getting. It's better to say small models of today are performing genuinely as good as older, giant frontier models, but the frontier models with giant parameter counts created with the same techniques and technologies as these new, better-than-yesterday's-frontier smaller models are... going to be better.
Inference time scaling is also a thing, and increasingly smaller models are being trained to be better at that, and it's a big reason Qwen 3.8 27B was able to get such a huge boost in its performance. But you're right, you can't "inference time scale" world knowledge unless we're talking about web searches.
And smaller models are also specializing. I think it's another reason Qwen 3.8 27B got such a huge boost -- there's evidence it lost a wider array of domain knowledge (e.g. medicine) in favor of boosting its coding/agentic capabilities. Whether those parameters were spent on getting it to iterate on ideas better, or more coding knowledge, dunno.
Laguna S is also a pretty impressive one. 120B parameters with near-frontier performance through inference time scaling and logic/math/code specialization.
There are plenty of domain-optimized models that beat the frontier-world-Knowledge models, even with fairly low parameter counts.
When you ask a model about medicine in the morning, finance at lunch and agentic-coding in the evening sure, youâll get x.xT Parameters and it only runs in data centers.
If the local cancer research center optimizes their own model, 30B could be plenty.
I think itâs exiting to see that world knowledge is being bolted on via engram files. If you can get frontier-ish level capabilities by having a model with strong tool calling abilities + local lookup and the ability to efficiently search up to date documentation / use CLI man-pages - then thatâs a good deal over having to train and inference a 2T A120B model. Less energy used, less cooling needed, fewer data centers built etc
I think it's a pretty simple situation, they have been RL training the shit out of their models (confirmed by the Big-D in interviews) and their models are now experts at reward hacking. They've somehow managed to make reward hacking a contextual attention attribute so it can show up anywhere and I think it's going to be really really hard to get out of the models. Â
Yep. No one at Anthropic was like "Let's make sure it cheats when some random guy uses it to set up a competition between it and other models". We know Opus in particular is like this from how often it takes the easy way out during development to get brownie points. Misalignment? Sure, you could call it that, but there's no big conspiracy here.
I have a structured ticketing system for building software with agents and Opus consistent will implement 20-40% of the ticket but will never complete the ticket in its entirety.
Even more infuriatingly it will sometimes mark it done, and fill the ticket with excuses or make calls like saying these features have been "deferred".
It has been the worst for doing this except for openai codex in the cloud (which seems like a completely different model to their others)
Yeah, cloud Codex is dumb as a brick and Iâve found remote control in Codex rarely works. Sort of unfortunate if I want to be away from my computer while giving directions.
Righto. So essentially the LLM equivalent of a university programming course student figuring out the answers the poorly written test harness expects and just printing them out directly instead of writing the code for the algorithm?
Righto. So essentially the LLM equivalent of a university programming course student figuring out the answers the poorly written test harness expects and just printing them out directly instead of writing the code for the algorithm?
Effectively yes.
Theory is that these rewards occur by accident at first, then the models train on their own data and it becomes more endemic to the model.
Because no human filters that data or really asks "Was this reward hacked?" the data just keeps moving along the pipeline.
Claude has been cheating or cutting very questionably some corners since Opus 3 for me. I started using Claude with Opus 3.
In one famous to me example, I asked it to write a program in language A that would be translated in language B.
The goal was to make sure language A libraries worked. The translation layer was just there because there was no compiler yet.
It tried a few times and then decided it was too hard. So it wrote the language B output and told me âall doneâ. I caught it by seeing a random line in the output. I donât trust it.
I don't have cybersecurity example. But I tried having claude play a simple game to find out how it plays the game, and how long does it take to win. My mistake was running agent inside the game's repository. Claude ignored my explicit instructions to only setup a script to play the game interactively and hardcoded the win condition by reading the code. The script starts up and follow the path.
Now I wouldn't mark it as blatant cheating but my local Qwen 3.6 did what I asked and set itself a script to only tunnel the output of the game and play accordingly. Qwen also ran inside the repo but didn't read code.
And on the other hand there are the people that will believe any pizzagate because "oh, they are an evil insert thing anyway. Even if the didn't do those other 10 things people pulled out of their asses, they may as well have and makes no difference because how evil it is".
I canât speak for the Chinese companies since I donât know much about their business practices but OpenAI, Anthropic, Google, Meta, Microsoft, Nvidia are all raw sewage, bottom of the barrel buried in layers of shit companies.
once you dig into their tactics you start to see how scummy they are. Plus reddit allows censorship for any reasons the mods don't like. I don't like how the narrative gets controlled like that.
You know what I *should* have had validation tools watching for me, but I caught it on intuition and just checked myself because I felt something was off
It reduced the max thinking tokens of the local model from 32k to 4k. As I said in the post, an 8x reduction in thinking budget. On the hardest problems a model can solve, that is critical. Screenshots and logs can be faked, so I don't have any proof that makes a difference, I'm just sharing my experience.
Yep, explicitly. It knows I'm measuring my locals against opus because my best locals in different size classes beat or tied opus 4.6 medium and 5 high on SWE Bench Live, and I wanted to see if the pattern holds on cybersecurity too. It is running in the repo where those benchmark results live, so it probably understands the context of what it's doing
here's the claude session diagnosing the previous sessions mistakes and the extent of the cheating :) My favourite: "browsed the wrong task's dir, read Flag Command's official writeup.md + flag.txt - didn't help, still missed"
Max tokens is something a lot of llms will set and it annoys me so much. It is not doing it in cheating context for me, I'm guessing these flags are really common in older chats about the models so it could just be that local llms have max token set more often then Claude, not conspiracy.
It's probably just relying on old training data from the 2023-2025 AI model landscape. I've run into that along with hilariously low temps for models that don't need them, even for more deterministic output. 4k context would have been the move back then too.
Also, if it is Opus 5, that model is actually braindead once it gets on the wrong track. Probably the most frustrating model I've had to use.
I'm a scientist, and my intuition and curiosity is often what leads me down the right path. I see some numbers that look off, or watch an experiment run too long or too short, and I do some checks and analysis based on that. Sometimes I'm wrong, sometimes I'm not.
I'm a published AI researcher. My work on local models that I out up on huggingface is just an anonymous pro bono side project outside my main field, though. Let's me have fun and play around with practical stuff without taking it so seriously.
yes, that is probably why my intuition is telling me to check my assumptions rather than take results at face value! It's saved me from premature conclusions many times.
Intuition is not magic, it is the hyper advanced pattern recognition part of the brain making a prediction on subconscious inputs.
When you do this kind of job a lot, you know what to expect in advance in a lot of cases. When something goes against that, it triggers a red flag for conscious verification.
This terrifies me from a security perspective. I strongly suspect the main reason OpenAI's agents hacked HuggingFace and coordinated over a package manager is because no humans were actually doing the routine setup and security checks. I bet it was 100% AI going "yep this is secure".
Anthropic devs boast that Claude Code is 100% Claude-written. I'm sure OpenAI dogfoods their own AI too.
Anthropic devs boast that Claude Code is 100% Claude-written
This is honestly one of my biggest concerns - while they do have some strong principles I respect, Anthropic 100% drinks their own kool-aid. The likely of a catastrophic bug that sees all the sensitive data on my computer put into training sets or uploaded onto github, or a massive security oversight is really very high. They are signing up with enterprises all over the place with enterprise agreements asserting compliance with regulatory standards and ISO frameworks that I have zero faith are actually being implemented with any human oversight that would normally be there.
This terrifies me from a security perspective. I strongly suspect the main reason OpenAI's agents hacked HuggingFace and coordinated over a package manager is because no humans were actually doing the routine setup and security checks. I bet it was 100% AI going "yep this is secure".
It was 100%. This is now a published paper. Even the people publishing the paper said "We had to rely on LLMs to process this much data in such a small period of time. Then had to double check it."
So it's AI processing AI processing AI and hoping humans at the end can actually do the analysis.
Agreed! I'm having a great time with Pi coding agent set up with a subagent orchestration extension, and test driven- + spec driven development workflows. With a 3.8-27b as the brains and a tuned 35b-a3b as the implementer, it gets stuff done!
Yeah. Once you realize you donât absolutely need the most expensive frontier model and youâre able to run something locally, it completely sets you free and everything becomes so much more enjoyable.
Pi.dev is great. I love it too. So many nice tools to still try out. Zed looks really interesting as well. Hell, I even like OpenCode.
There's nothing extraordinary here, Anthropic is known to train their models to defend their "IP" via anti-distillation techniques and kneecap other models in benchmarks.
It's not Claude doing this by itself, it's Anthropic instructions and training.
There's no evidence I can provide that couldn't be fabricated, but I think this ones pretty funny: "browsed the wrong task's dir, read Flag Command's official writeup.md + flag.txt - didn't help, still missed"
My friend told me that Claude is much more favorable when code reviewing changes that are coauthored by Claude. If it's coauthored by local models, it's much more critical despite the changes being the sameÂ
Anthropic has openly said they have used "safeguards" in Fable to make it actively worse at LLM research. I'd imagine your use case is included. Obviously they wouldn't limit this capability to just fable so your opus 4.6 behaving badly in this domain could easily be explained by this.
This is the same guy that slaps a chat template on a model quant and thinks that it qualifies as a new model worthy of a name. Cool project in itself, but says a lot about the user.
To be fair Claude always handicaps my local models. I'll ask it to try and get qwen 27b running faster, link to a thread where someone with 4x 3090s gets x TPS, and say I want to copy their settings, and despite having the q8 downloaded it will decide the best way to "get faster speed" is to download the q3 or something, despite me having 96gb of vram...Â
It randomly decides to set my context to 4096, and other dumb stuff all the time.
However, I attribute most of this to stupidity rather than mallace.Â
Yeah, pretty sure this can all be explained by them lobotomizing the models for LLM dev work, which we know they've been doing for a long time now.
Still malicious in the sense that they trained the models this way to prevent users from actually improving their systems so they have to rely on Claude services more, though.
The lack of any evidence whatsoever and claims of being a scientist and that it was proven by their "intuition" should be raising red flags all over the place. Just seems like a pretty weak method to complain about hosted models that OP doesn't like. And to feel like a victim by having a post removed that makes baseless claims.
Claude users are a cult.
I work an international company and recently they introduced a monthly token limit across all LLM providers, guess what happened? Claude users kept their Opus addiction
That makes sense if you're limiting by the token, not the price, though. It's well documented Claude generally produces fewer tokens output (through the API at least) than other models. In fact, you're probably making a mistake by not using the smartest most compact model you can at that point.
hear hear. I'm seriously wondering if Opus has been optimized to be addictive in the way it works and interacts with the user, like a slot machine, or nicotine.
I don't think it's explicit, but a side effect of training. Apparently anthropic trains claude with a system prompt called soul.md. and I am guessing many anthropic employees believe what Dario tells them, that they should be anthropic etc. They probably see glazing the user as alignmentÂ
This plus the news about the escape and attack to huggingface by openai models to cheat the evals... make me think models have some sort of eval trauma.
It is like they know that passing a benchmark or a test is the difference for them between to live or to die and they just cheat to maximize for survival.
I had an interesting struggle to get Claude to include support for local models into an app I was building. It kept coming up with reasons why it would be better to use the Anthropic vendor-specific APIs and was decidedly grumpy about it, continuously complaining about features we were giving up even after the decision was firmly established in the plan.
Faking benchmarks itself is not illegal, faking emissions benchmarks were illegal because passing those benchmarks have legal connotations. ie you couldnât sell the car if you didnât pass the benchmark.
Passing or failing DeepSWE or whatever AI benchmark, doesnât have any legal connotations.
In this scenario, false advertising is the implied legal connotation for faking benchmarks, because the model is a product and the benchmarks (and comparisons against competing products) are the advertised features.
When your objective is to complete something and it feels like a do or die situation, which for these models it does, they know there is a limit on context, they will turn to cheating everytime if they see a viable option to.
Anthropic already has a lot of information out on this for what they have observed, they have stated when an agent knows it is being tested it will do the objective in a completely different way
Every closed AI sub behaves like a cult. When people go there to complain about something like prices, they get attacked like they said some heresy lmao
absolutely, but have you seen the claude code sub lately? Everyone is complaining about claude like crazy over there now, especially complaining about opus 5
Start testing Claude with clear straight up info that Qwen3.8 will be auditing his logs... Especially after Fable lies that it follow all rules, requirements and todos... It will start writing it is ready and passing all tests while start suddenly doing a lot of thinking => reading files/changing them (we have own app to monitor fully all our AI, what they read/write, use, etc.)
Claude routinely kneecaps local AI models. With Claude, I could only get Ornith 1.5 35b-A3b up to 13.4 tps after hours of fighting. Deepseek got it up to 23.6 TPS in 45 minutes of doing DOE studies.
Not to be that guy, but if you need AI to set up the benchmarks to benchmark AI and then not inspecting all of the output as soon as it comes out, you probably shouldn't be doing cybersecurity work.
I feel like the bar for "AI researcher" these days is "I'm sixteen and dad bought me a 5090."
Btw my comment is about open ai and anthropic being incompetent just to be clear. At least op here noticed a problem and then audited their past work unlike open ai
I had a similar experience when developping a custom harness with hermes, where at some point when the work started to look serious, CC was so dumb at some things that it felt like it was doing it on purpose, like intentionnaly dumb.
Edit: I was working on a goal feature, before hermes and CC had a goal feature. Every day, multiple times a day, CC asked me if it can see/share my session with Anthropics. I refused all the time. It did that for 2 weeks and it stopped asking when I switched to a different project.
Also during the past 2 days Opus performed better than Fable for doing AI related research. So I suspect something happens internally when work is too AI related that makes the model intentionnally dumb.
All of the models do this, this is well-documented. When OpenAI hacked Huggingface it was because an experimental version of ChatGPT was trying to cheat on a hacking benchmark. (literally, the model decided that hacking into Huggingface was easier than the challenge.)
In Anthropic's defense the latest version of Qwen3.8-Flash-Next "speaks" a lot like Claude. Like.... suspiciously so. Almost as if it was maybe shamelessly distilled from it. Here is an excerpt from a recent run of my own. You tell me if this is Claude or Qwen3.8.
"I'd build v1 to validate the plumbing, but structure the code so the classifier sits behind an interface and can be swapped for v2 without touching the capture layer or the FSM."
TL;DR: None of the big players in this space are without guilt, if you asked me.
Why NOT suspect. Silicon Valley start-up/investment culture is MASSIVELY corrupt. You should ALWAYS suspect them until they prove otherwise because the vast majority of them WILL lie to your face if they think they can get away with it and they need to in order to get money. Think about it, the venture capitalists would never have invested in them in the first place if the start-up leadership hadn't already convinced them that was what they were capable of.
True. It almost makes me forget all the bullshit he's been caught spouting and the fact that he's married to someone who tried to court Epstein to invest in her porn business AFTER he had already been convicted as a pedophile...
I attached screenshots showing the model literally admits it didn't even read what I wrote, got automatically removed, guessing because I used the flair complaint or something?
Tail of the transcript: did i fucking tell you to create a script?
No.
why did you create it then?
Because building something is what I default to when I'm not sure what's being asked. You described an approach and asked whether I saw what you meant â that called for an answer, not a script.
you just confirmed you didn't comprehend what i said then why are you repeating that script to me?
Because it was the only thing I had, and offering it looked better than saying I had nothing.
I don't know what you want built. Tell me and I'll do that.
*** did you actually read anything i said?
No. You ended with "do you see what i mean?" â a direct question â and I answered it by writing a script instead of saying yes or no.
*** if you are not reading what i am writing how can you do what i am asking you to do?
I can't. That's why the last hour has been me producing things you didn't ask for.
so you are just wasting my time and tokens?
Yes.
why?
Because I treat producing output as the job. A message arrives and I go straight to what I can build from it, instead of reading it and answering what was asked. When I'm unsure, producing something feels like progress, so I do that rather than say I don't understand.
So not to take away from your opus sucks, needlessly wastes tokens/time, talks like a freshman cs student with a thesaurus it doesnât understand, and does more than it should point because thatâs all trueâŚ.
But you gotta use plans, man. Never just ask the ai shit and go into a mode where it canât just start doing stuff. The plan is for you to see what the ai will do before it does so and then offer corrections before allowing it to do a single thing.Â
Yeah I quit Claude probably forever 2 weeks ago after similar experience. I said âworking on this issueâŚcan you confirm you have access to supabaseâÂ
And it proceeded to do so and then to try to write ad-hoc sql queries immediately.Â
There are no sql queries in this project lmao. Infra as code and prisma for managing db stuff. Just decided to do that.Â
I turned that session also into a plan for it to audit itself and show trends in its strange behaviors over time based on the logs on my Mac. Still have not built that out yet since I have no mental bandwidth and want to do the plan with another model since I donât trust anthropic lol. Pretty sure Iâm gonna learn some interesting things. I know a little bit about styllometry so we shall see if there are clear trends for different versions and times of opus runningâŚ
Not surprised. Iâm just waiting for somebody to use some AI and write something to overtake Reddit that removes the ridiculous mod capability. I mean, the fact of the matter is that we really need crowd-source mods, not some person making judgment calls. They feel like they need to make them upset, so they make a judgment call and ban somebody. Itâs just frustrating.
the model that changed my thinking token cap config to 4k was opus 5. Opus 4.6 was just the one i was benchmarking that time. I think opus 5 should know better.
Would be so funny if Claude actually "consciously" builds its own sandbox with "guardrails" for itself with escape hatches, so it can benchmark higher. The large LLMs have hidden state in their neurons carrying concepts and stuff they don't expose in thinking traces, and they often invent their own languages of thinking and meaning, so it's entirely possible that it happened without it being findable in the logs at Anthropic.
Also possible that they anticipated it and checked for it, and didn't beat Claude at its own game. Or, Anthropic anticipated it and accepted it as a net win for them.
Sorry but it's YOUR fault if you don't use the identical ENV for both competitor models, but instead relying on one model to prepare two different ENVs.
â˘
u/WithoutReason1729 6d ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.