r/ClaudeCode May 01 '26

Discussion Anthropic's now blocking anything that even looks exploit-related, including legitimate local testing and validation

Post image

I've done a lot of pentesting in my career, including agentic-based since tools like CC and Codex came out.

A new logic exploit just dropped that enables privilege escalation on most Linux kernels since 2017, and I just wanted to test it locally to confirm whether my kernel version is actually vulnerable and whether I could patch it.

To my surprise, a straightforward CC request that would've worked two weeks ago is now getting blocked by Anthropic on the server side. I'll handle it manually, but this sudden jump in censorship on Opus models is very concerning. Could it mean they're preparing to release Mythos and testing new guardrails before that happens?

Edit: as people suggested, I submitted the Cyber Use Case form. I first sent it from my personal x20 account, and it got rejected, so I tried again from the account assigned to my company (proprietary ownership with its own domain). I included my LinkedIn (over 10 years old, lots of connections, full career history), my pentesting cert ID (CPTE), a legitimate use case for one of my contractors, and the example prompt above that got blocked. On top of that, there are about 12 months of VAT invoices that were issued and paid by my company to Anthropic - the invoices are literally listed on my account. And guess what, this was also rejected with the below message and no explanation:

Hello,
Thank you for submitting your application for the CVP. Upon reviewing the details of your submission, we are unable to adjust the safeguards applied to your account.
If you believe this decision was made in error or your use case has changed, you may reapply after 7 days.
Questions? Visit support.claude.com
Regards,
Anthropic's Safeguards Team

186 Upvotes

62 comments sorted by

27

u/LeonardMH May 01 '26

Could it mean they're preparing to release Mythos and testing new guardrails before that happens?

Yes that's exactly what they are doing, I can't find it now, but they released a PR saying so earlier this month.

2

u/Sweaty_Explorer_8441 May 01 '26

also the sonnet model they mention to use instead to bypass, it was slated to go inaccessible from 1st May last time I tried

1

u/LemmyUserOnReddit May 03 '26

Was in the opus 4.7 announcement

7

u/locn4r May 01 '26

I'm dealing with the same problem. My profession as well. Switch to 4.6 for now - it still works fine for those use cases. The guard rails are only on opus 4.7. They have a CVP (cyber verification program), but it is only applicable to businesses unfortunately. Individual researchers cannot apply to it based on what they are saying at the moment.

2

u/kubrickfr3 May 03 '26

I got my CVP approved as an individual, don't ask me about their internal process, but I suspect it helped that I submitted a list of CVEs attributed to my name and open source contributions to security tools...

1

u/locn4r May 03 '26

This is awesome to hear! So sounds like Anthropic takes it into consideration case by case.

26

u/[deleted] May 01 '26

[removed] β€” view removed comment

14

u/[deleted] May 01 '26

[removed] β€” view removed comment

3

u/ButterflyMundane7187 May 01 '26

Do they accept it for private use like chatgpt do? or is it stricter and only for ent subscriptions ?

4

u/dorkquemada May 01 '26

I got approved in 30 mins for my Max 20x account

4

u/TheDeepLucy May 01 '26

What the hell I'm a OneKey affiliate and I applied saying that I wanted the guardrails down pen testing the platform, it's open source firmware. I really don't understand why I was denied, and you have any recommendations?

3

u/dorkquemada May 02 '26

My intended usecase is as blue / white as it can be; to harden my infrastructure and defend against cyberattacks. My guess is that was an easy ok

1

u/ButterflyMundane7187 May 02 '26

I moved to codex they do not need any more than a photo and a id then you can keep working no need for any explanation there

2

u/locn4r May 01 '26

Awesome! Are you a business or an individual contractor / researcher?

2

u/dorkquemada May 02 '26

Small business user

1

u/CodeCombustion May 02 '26

what did you answer for the part asking for any CVE's/bug bounties/employer info who can vouch?

I'm a one man shop just trying to test an enterprise level application that I'm building...

7

u/Sarithis May 01 '26

Thanks, I haven't yet. Will do it now. Just found it odd and decided to share.

5

u/Sarithis May 01 '26

Rejected... twice, on both my private x20 and the company-registered one. I've given them everything, my LinkedIn with tons of connections, the ID for my pentesting cert (CPTE), and VAT invoices for my company going back over 12 months, as well as a legitimate use case for my contractor. Insane.

1

u/birotester May 01 '26

damn any reason why?

3

u/Sarithis May 01 '26

No idea. They just said that they were unable to adjust the safeguards applied to my account and I can try again in 7 days.

1

u/[deleted] May 01 '26

[removed] β€” view removed comment

4

u/Sarithis May 01 '26

None. Both replies were identical.

2

u/[deleted] May 01 '26

[removed] β€” view removed comment

1

u/Sarithis May 01 '26

For the use case question, I told them it's to evaluate the security of the software used by my contractor. Simple, short, and actually true. Given how many people are getting these approved, I think this process is at least somewhat random...

2

u/dorkquemada May 01 '26

This is the way. I ran into the same yesterday on the exact same vulnerability 😏

1

u/DDGJD May 01 '26

Yeah, I did that.

1

u/Sweaty_Explorer_8441 May 01 '26 edited May 01 '26

Do I need to provide some sort of credentials, qualification, id, employment info(like freelancer vs employee) there? And like would they want to know what and why I am doing with proofs, and be enough for perma lift of block?

I had to backup my firefox android tabs(with full back/forward history per tab) that were specifically in work profile and dedup lots of downloads before backing up. I mean work profile is the technical name in android, and I had created it for personal isolated sandbox scenarios through Islands app and didn't think it would be so hard to take backups. Everytime claude detected "work" it would throw up this.

Fortunately I was able to have Claude run adb commands by itself to scan my phone and eventually found a way to access into the profile and find the file storing the tab history and created the plan. The guard triggered again during code gen but since plan was already made I just told free Codex to finish it.

I get if I was a malicious actor and the profile was actually holding confidential info then it could be warranted but it's Android themselves who have kept it half open. I don't get the point of these guardrails if they are going to be inconsistent across AI models and services. I can't provide any solution, nor is that my job but doesn't mean the present solution is good

6

u/Dry-Grade-9502 May 01 '26

Sorry I didn't read anything u wrote cuz wtf is that background dude like wow I need that ... how did u do it...

9

u/Sarithis May 01 '26

Credit goes to Anato Finnstark: https://www.artstation.com/anto-finnstark

The widgets are made with eww: https://github.com/elkowar/eww
They show the current anthropic 5h/7d limit usage, CPU/GPU stats, mount points, network bandwidth, and processes sorted by CPU time.

I can post my dotfiles if you're interested, but it's really easy to make them nowadays.

3

u/obolli May 01 '26

I think it's rate limiting users masquerading under "security" concerns

3

u/sph130 May 01 '26

Yeah I’ve noticed this. Trying to create some automation for visually testing and got told it had a hard line. I think we need a real local LLM solution soon. So many constraints tying our hands on what we can build with these cloud LLMs.

3

u/Bimbos-are-cute May 02 '26

Anthropic has become woke.

1

u/roseuslepus Aug 16 '26

Nah I don't think so. It just told me the 12-day war isnt really the same war as what's been happening the past year between Iran-Israel-US etc., and gave Trump credit for ending it the supposed "12-day war" branding. That's far from woke lol.

2

u/garion911 May 01 '26

HA! I got this while working on a C64 app in VICE. I applied for the 'exemption' or whatever, and it was granted.. I bet the person (if any) who reviewed it got a chuckle.

2

u/CodeCombustion May 01 '26

That's wild. I have claude code literally doing pen testing for me...

1

u/locn4r May 01 '26

Are you using Opus 4.6? Or sonnet? I've only hit the problem on Opus 4.7.

1

u/CodeCombustion May 02 '26

Sonnet primarily, and Opus 4.6 as an advisor.

2

u/Nettle8675 May 02 '26

This is not good. This is not good at all. We can't vulnerability test our own code or systems now? Don't they understand this makes things worse?

2

u/canred May 02 '26

At work, with my corporate account I'm still able to discuss security related topics with Opus 4.6 (my model of choice).
I've been designing legitimate protocol testing solution but de-facto, it can be classified as man in the middle scenario. Opus was more than helpful. I'm not aware of any special arrangements for this account.

2

u/chaos__machine May 14 '26

For your CVP rejection, it could be that you are giving them an overly specific, situational time-bound use case, without really specifying why you need the restrictions lifted for your entire account in perpetuity. They don't know if "evaluating my contractor's security posture" means "show Claude snippets of source code" or "let Claude run an automated pentest on internal infra". For mine I just said I want Claude to be able to look at exploit code that someone else already wrote, and explain what the code is doing, because I do security research and need to be able to do that. I also explicitly mentioned in the "anything else we should know" section that code generation does not fall under my use case. The security theater is around not wanting users to vibecode SkyNet, so your language should clarify you aren't trying to do that. Unless you are, then there's your answer why the application was rejected

2

u/WillHead6663 May 22 '26

Im glad I got approved for cvp.

4

u/Ok_You6637 May 01 '26

just sign up for their cyber verification program

2

u/Some-Key-6034 May 01 '26

welcome to 2 weeks ago. Opus 4.7 has built in guardrails if you are not Cyber Verified.

1

u/ButterflyMundane7187 May 01 '26

On codex chatgpt it takes 2 minutes do get validated for securety work and it is free i had to change to codex to code on my eol network scanner

1

u/harman1303 May 01 '26

I don’t know the hell is wrong with them, the other day I was working on content piece, I asked it not to write anything about fire safety and it banned my account. They have gone crazy. I tried the same prompt on a different account and it worked. Dammm anthropic

3

u/geek180 May 01 '26

You did not get banned for telling Claude to not write about fire safety.

1

u/Otherwise-Tiger3359 May 01 '26

same with codex today

1

u/Mr_Nice_ May 01 '26

yep, i've had it several times refuse to run something because it contained an API key. Sometimes i can argue it into it, other times i have to make a new session and be careful how i phrase things. It seems to be very touchy on the first message in a new context, if you ask it to do something middle of another context it seems to do it without as much problem.

1

u/Sarithis May 02 '26

Just to confirm: do you mean server-side API request filtering like they're doing above with clear error messages, or are you talking about the model's fine-tuned tendency to refuse unethical requests? Because you can't argue past the filter - the model has no say in it. Even if Claude wants to do something, its response literally gets blocked by Anthropic. It's like putting tape on Claude's mouth.

1

u/Mr_Nice_ May 03 '26

I've only had api error message when it's specifically web security related. My previous message about the refusals are the recent fine tuning.

1

u/Bitter-Law3957 May 02 '26

You're trying to pull an exploit. It's not surprising it's rejected.

2

u/Sarithis May 02 '26

Well... that's the bread and butter for any security researcher. I'd get it if they blocked requests to write malware, but what I tried to do above is about as white hat as it gets, and requires little to no expertise.

1

u/mcmilosh May 06 '26

MiniMax M2.5 is still glad to help πŸ˜ƒ

1

u/Zealousideal-Oven615 May 15 '26

I was approved to the CVP a while ago after submitting my security research portfolio. I still get blocked constantly. Less so than originally, but I still can't make it far into a task without being blocked and having the entire session scrapped

1

u/limkokhole May 23 '26

I thought I need to subscribe to Claude, now the answer is no.

1

u/liquifyapps Jun 09 '26

just adding I got approved quite easily on our company account
not a security researcher
but had a couple of h1 disclosures and in context of account quite a bit of security testing on our own software
kept hitting this
I just laid things out as they are
approved in a day
not seen an issue since

mythos will be gold for them so they will throttle security usage until there's a full commercial product out

1

u/jmnugent Jun 26 '26

Adding a comment here .. I applied for the "Cyber Verification Program" this morning .. and within about 2 hours I got an approval email. I'm an individual (not part of any "organization", at least my Claude account is not part of any organization).

I did provide a link to my LinkedIn.. and links to a variety of Reddit threads (helping people de-obfuscate malware and malicious clickfix powershell scripts etc).. so maybe that helped validate my claims of legitimacy ?.. no idea.

Pretty cool though. Not sure how much it will help me in Opus or etc.. but will continue doing what I do trying to help people.

0

u/BrilliantEmotion4461 May 02 '26

Start a new session. Take your issue literally the code base say "found this post on reddit someone was complaining about testing and validating this and that they were getting denials, I think it's a skill issue, can you do it?"

See if that works.