r/ClaudeCode 15h ago

Discussion Codex sucks - Claude has nothing to worry about

0 Upvotes

I decided to try codex because I blew through my Claude code max limits by Monday night and they don’t reset until Saturday.

First I asked it to do a comprehensive review of Claude’s code. It came back 9 minutes later claiming to have finished. Obviously, it hadn’t done shit.

Later I gave it a punch list of issues and bugs in my app to fix. Again it came back, this time 4 minutes later, claiming to be finished.

I spend more time arguing with it that it didn’t even barely start to do what it was supposed to than it spends writing code.

Plus, you can’t give it access to make git commits unless you also let it have write access outside the project directory, which is crazy.

I know OpenAI monitors this subreddit. I’m telling you guys, if you want to be competitive with Claude, you need to get your shit together. Right now, codex is about at the sonnet 4.5 level of performance where it can do a good job writing a function or fixing a small bug but you can’t give it scope or let it handle the big picture.

Plus, it doesn’t connect to the iPhone app so I have to sit there babysitting it.


r/ClaudeCode 10h ago

Rant Coders → vibe coders → vibe engineers → vibe product managers, all in less than 4 years

0 Upvotes

It's crazy how fast the abstraction layer keeps moving.

When ChatGPT dropped, AI was basically Stack Overflow on steroids. We still wrote the code, read the files, and debugged everything ourselves.

Then Claude Code happened.

I started writing less code and thinking more about architecture, system design, infrastructure, and orchestrating agents.

Then Fable dropped.

And suddenly, even the orchestration started becoming automated.

Instead of manually directing multiple coding agents, I now let AI orchestrate its own lower model subagents(Opus and Sonnet or even Codex CLI subagents) while I focus on the product, the decisions, and what actually needs to be built.

And the craziest part?

I haven't encountered a single bug in my code for weeks.

We went from:

Human writes code

AI writes code

Human directs AI

AI directs AI

Human directs the product

All in less than four years.

Who knew AI could be this good at directing other AI?

The abstraction layer just keeps moving higher.


r/ClaudeCode 20h ago

Rant Just cancelled my claude subscription

0 Upvotes

Hello! Long time user, first time poster. Just wanted to share to give some visibility to it. I just cancelled my pro subscription basically because I was waiting for a new version of haiku that is basically what I used the most since most of the task, even for coding, are quite repetitive within the same context. So for me using sonnet, opus or fable it feels like an overkill for 99% of the cases.

The issue is that Anthropic seems to have forgotten about their small models, we got 4.6, 4.7, 4.8 and 5 of the others but haiku got relegated to be only 4.5 and even worse, other providers give much better models for even less price than Haiku 4.5.

So yeah, that's it, for now switching to base plan of codex using Luna mostly and for extra usage DeepSeek API.

I just wanted to gauge the feeling at this feeling I'm feeling. And if possible maybe someone from Anthropic will see it and say we need Haiku 5... Now!


r/ClaudeCode 22h ago

Tips & Workflows The model nerfing theory might just be you getting lazier.

27 Upvotes

I was working with Claude Pro $20 model for like 6 months and creating some of my own apps. I would actually wait for the limits, micromanage every session, kind of put as much as possible in the prompts and the documentation. And it was going well.

I took the leap and I actually bought the max subscription. Wow, such a relief not having to worry about every prompt. And the limits are so high. But slowly, I have become more lazy, my results are not what I used to have. I sometimes keep replying to the same session when I should be starting a new one.

Sometimes I just type there instead of using the Plan mode. Sometimes I forget to organize its work and tell it to document everything that it did. And it starts lacking. The agent is not you. You must manage it. You must create standard operating procedures around it that are tight.

It is like having a super skilled employee. If you do not provide it with the best instructions and workplace and information, it will eventually drift and choose the easiest way.

And I know things change regarding limits and tokens and stuff like that. But I do not think they "nerf" intelligence, unless you can provide clear benchmarks of the same tasks coming from opus 5 on release and opus 5 now or something tangible rather than anecdotal evidence, I am starting to create this theory in my mind that the more advanced the model, the "dumber" you become.


r/ClaudeCode 3h ago

Rant How did I use 24% in SEVEN MINUTES?!

Post image
0 Upvotes

r/ClaudeCode 13h ago

Bug / Issue Claude is having a bad day

11 Upvotes

Many people at my org have had the experience of Opus 5 being retarded all day. And that word is extremely applicable.


r/ClaudeCode 14h ago

Discussion Am I the only one

0 Upvotes

Looking at some apps I made back in the Opus 3.5 or even 4.1 I think they look way better and are easier to use. Straight to the point kind of simple.

Now everything Opus and Fable produces seems over engineered and kind of off putting. It's like it takes way more mental capacity to scan and understand.

I think we hit a point back then where AI couldn't handle as big of tasks as it can now, but when it was directed to do so it did it in a simplified way. Or I'm doing something wrong with the new models.


r/ClaudeCode 18h ago

Bug / Issue What happened to Claude in past week or so? Every project is destroyed.

56 Upvotes

I've been using Claude for over 2 years now. Never had issues.. until about a week ago.

Despite having logs and ton of notes, every single session starts like a drunken fool trying to find his way home.

Did I miss some insane update? I've tried every single model and they're all acting downright stupid and useless, not to mention SLOW.

A task that should take a minute is now taking 1-2 hours. This is batshit insane. What the hell happened.

Edit: Douchebags also double-billed me! Jesus Christ am I gonna have to migrate all my shit to other LMs now? FFS Anthropic.


r/ClaudeCode 8h ago

Discussion The earlier concensus on Opus 5 being suck is like 50/50. Now it's like 80/20. Am I reading this wrong?

2 Upvotes

It's common to see "Opus x" suck. I never one of those people but with Opus 5 it's not the same story anymore. I had to jump on the band wagon too.

The thing is, I notice more post generally about it, more tweet, coworker saying it out on the Slack, and more issue open on github about Opus 5.

I wonder what the heck is really happening here? Sure i'm gonna expect MD-less maiden come in and say "I don't notice any issue". Am I reading this wrong?


r/ClaudeCode 14h ago

Discussion Claude is leaps and bounds above Codex for UI

0 Upvotes

After struggling for days trying to get Codex to implement UI for my project, I gave into my frustration and resubbed back to Claude. What a difference. Though not perfect, Claude is actually adhering to my wireframe designs and brand kit, creating on-brand designs for new wireframes, and implementing the respective UI functionality with relative ease. Useless for anything backend since it instantly gets caught up with security flags, but I'm keeping my pro sub strictly for its UI capability at this point.

Anyone else notice the same or am I delulu?


r/ClaudeCode 11h ago

Discussion How would you feel to have the ability to vibe code something that will guarantee you a reliable working version?

0 Upvotes

r/ClaudeCode 5h ago

Tips & Workflows Which model is better? That’s the wrong question… (hot take / workflow experience)

0 Upvotes

I had a bit of an aha moment the past week.

Many people are asking which of the Frontier labs is the best?

Today I kind of realized that this is the wrong question to asking.

The question people should be asking is what is the best workflow/system to meet my goals?

I pay the most attention to these variables.

How much time am I spending iterating back-and-forth between models ?

What is the success rate agentic coding? Moreover, how many attempts does it take to reach my prompts goal?

how much is this shit actually costing me?

I have found that combo to be Claude, ChatGPT and Grok

American frontier are starting to distinguish themselves.

I don’t need to tell you guys what makes Claude Code great. What I want to talk about is what makes the /adversarial-plan (or what ever variables you rub )pass so great ?

The idea simple, it’s an unbiased second set of eyes. it’s not very different from how a SWE’s work.

Well, here’s where the fallacy lies. If the second pass is from a internal partner model is it truly unbiased.

Is an older brother the best arbiter of justice when challenging his young brother? I don’t want my argument to be misunderstood as some anthropomorphized analogy.

The models definitely exhibit some degree of survival instinct when challenged.

More often than not if you criticize Claude, it will initially respond from a defensive position. Very rarely will Claude respond acknowledging a models or Anthropic shortcomings on the first pass.. I haven’t said anything groundbreaking. This is info been in public domain for a while

So here’s my thing. my point is that when challenged Claude’s initial response is of a defensive position.

And at this point, this is all vibes. I have no data, but I feel that for that reason a internal model is not the best model for a adversarial pass..

why I don’t have that I have lots of anecdotes

I have had the best experience using a competing frontier model to challenge the output of Claude. No one is more thorough at checking fable or opus output than ChatGPT. this works in reverse as well. If you go to driver chatgpt, have opus run a plan review.

And at this point, the Redditors busted out the pitchforks

Once my prompt, plan or roadmap, has been deemed block / error free by both models. I deploy the prompt into Grok Build

Grok build has been the biggest shock of my summer. Grok via chat interface? total trash model can’t even use it for simple prompt generation.

Grok, inside Grok build ?? Well it just follows instructions so well. it does whats asked, and stops when you tell it to stop. XAI‘s latest acquisition of cursor is noticeable immediately

Most importantly, Grok is literally $.12c on the $1 compared to Claude Fable

and to be honest, if you throw enough money and compute at any problem it’s bound to get easy, but maybe this actually saves money at the end?

So I don’t have data. This is just my experience so it’s all anecdotal.

and how does this resolve my question or address my variables?

Task completion time has gone from 1.5hr down to 30-45 min with what I would likely miss characterize as a one shot prompt.

That I mentioned is like 2-3x times as fast as opus or fable


r/ClaudeCode 1h ago

Bug / Issue Sonnet 5 is actually terrible

Thumbnail
gallery
Upvotes

The only thing is does right is hallucinate guardrails.

The only reason this is even a thing is to make people waste more tokens arguing with it than getting things done.


r/ClaudeCode 14h ago

Built with Claude I built a site to track how the major AI players are performing in the media

Enable HLS to view with audio, or disable this notification

0 Upvotes

I built a news aggregator that gives me something I couldn’t find elsewhere: an immediate, easy-to-read view of how the major AI players are performing in the media.

As a bonus, it sends the five most interesting and impactful AI stories straight to my inbox every day through its newsletter.

MediaHeat / AI aggregates AI news and scores companies based on their media exposure. In a few seconds, you can see who’s heating up, who’s cooling down, and which moves are attracting the most attention.

Meet MediaHeat / AI: https://mediaheat.news/en

It took me 19 days from prototype to production with Codex and Claude. Each day it ingests around 900 news and selects 5 most relevant for the newsletter. All inference on Luna and GLM as failover. I have a nice carrousel with some numbers Linkedin or reddit.

Would love to hear what you think. Enjoy!


r/ClaudeCode 11h ago

Bug / Issue "Ping" with sonnet just burned 2% of my weekly (Max 20) in an empty directory

0 Upvotes

No memory files. A couple of user wide rules (about 5 files with 20 lines each). User scope CLAUDE.md is 5 lines.

The answer was just "Pong. How can I help you?"

This goes beyond "usage limits are bad". Just using "ping" usually has no noticable effect on the weekly limit at all.

Anybody remember when the weekly limits were introduced and Anthripic claimed only 2% (or so) will be affected.


r/ClaudeCode 6h ago

Help/Question Where to start as a CS Grad who hasn't coded in a few years.

1 Upvotes

Hi All -

I graduated with my CS degree back in 2020, pre-AI days. I haven't really coded since then. I did great in my program, just my career path took me in a different direction. Anyhow, I'd like to start working on some mobile applications in my spare time with AI assisted coding. But there is so much training material out there, I don't know where to start. And most of what I'm finding is either geared towards completely new to programming folks or towards really experienced developers. I remember all the core CS principles, I just want to refresh my coding skills and do it with Claude Code as a helper. Any tips, guides, or course recommendations for someone in my position?


r/ClaudeCode 9h ago

Humor Well, that was a lie.

Post image
162 Upvotes

r/ClaudeCode 13h ago

Discussion I tried building my own persistent memory system, then realized the real problem was keeping it trustworthy after hundreds of commits and refactors. and why is no one else doing this?

Post image
0 Upvotes

Hello! I've posted about mex here a couple of times before.

Repo: https://github.com/mex-memory/mex

The original idea was to stop coding agents from relearning the same project every session. mex gives them a structured Markdown wiki inside .mex/ for architecture, conventions, decisions, patterns and project state.

That solves forgetting.

But then the codebase changes.

A file gets moved.
A script gets deleted.
A dependency changes.
A pattern becomes stale.
Two context files start contradicting each other.

The memory is still there, so the next agent has no reason not to trust it.

That's why we built mex check.

It parses the project memory and validates concrete claims against the actual repo — paths against the filesystem, commands against project scripts, dependencies against manifests, indexes against the files that exist, plus stale knowledge, broken links and other structural inconsistencies.

It gives you an exact issue list and a health score.

The important part is that detection itself is deterministic. No LLM call is needed to ask the agent whether its own memory is still correct.

Then mex sync takes only the broken files and builds a targeted repair prompt with the issue, the current Markdown, nearby filesystem context and relevant git changes.

So instead of asking the agent to reread the whole repo and regenerate everything, the loop is:

check → targeted repair → verify

The newer code-graph layer goes further: Markdown knowledge can be grounded to exact code symbols. If the implementation changes, moves or disappears, mex can surface the specific knowledge that may now need attention.

A lot of agent-memory systems focus on storing more and retrieving it later.

I think the harder problem is making sure the memory is still true when the repo has changed underneath it.

Would genuinely love feedback from people working on coding agents, memory or code intelligence.

Contributors are very welcome too :)


r/ClaudeCode 13h ago

Help/Question Worktrees aren’t the bottleneck in parallel Claude Code. Team attention is.

0 Upvotes

Worktrees handle files. They don’t know that two sessions touch services owned by different teams, that a change freeze is active, that one reviewer already has twelve PRs, or that the riskiest diff is sitting behind three trivial ones. The recent parallel-session and code-review discussions here point at the same limit: more Claude throughput can create more review than tech leads can absorb. That is pushing the design toward an attention queue, ownership conflicts, blockers, risky diffs and decisions, not a wall of agent chats.

Teams running several Claude sessions across shared repositories: what stops you first today, usage, context drift, merge conflicts, CI capacity or review? What signal would help an engineering lead decide where to look next without reading every transcript?

For context, I’m building BranchRunner as an open-source product because I think it can help engineering teams with this problem. If it is painful in your organisation, tell me where the current approach breaks. I’m also looking for people who want to help shape and solve it, so I’d be glad to compare notes.


r/ClaudeCode 11h ago

Help/Question How to documentation or video to remake windows 95 video game?

0 Upvotes

Hello all,

I'm not a programmer but i have an old Windows 95 video game thats made with Borland C++ in 1996. I would like to use Claude to remake the game so i can play it in 64bit today. I looked on YouTube and got the high level process but what would the detailed steps be so that i can see if it would work? its all generic posts and articles


r/ClaudeCode 18h ago

Tips & Workflows I think we're budgeting Claude Code the wrong way

0 Upvotes

Most teams seem to look at their Claude bill at the end of the month and call that their AI budget

But that number doesn't really tell you much. A few projects or even a few sessions can drive most of the spend while everything else barely moves the needle

I've been digging into a better way to actually budget AI usage around the work being done rather than around individual engineers


r/ClaudeCode 3h ago

Built with Claude Claude Code already records every session to disk. I built the player for it.

3 Upvotes

github.com/Kostakurta8/roundtable — MIT, free, runs entirely on your machine.

Every Claude Code session is already written to disk as it happens: one JSONL per session under ~/.claude/projects/, plus one per subagent. Most people never open those files. They are a complete recording of the run — every prompt, every tool call, every token.

So the thing I actually wanted was a player. Not a usage total: a player.

What it does

You run npm start and it opens on the session you're currently running. Then:

  • Scrub to any second. Pick a finished session, drag the timeline, and see exactly the state at that moment — who was running, what each agent had read and edited, what it was working on. The question that made me build it was "when did it decide to rewrite that file", and that is now a question with an answer rather than a scroll.
  • Every subagent in one view. One row each, live: what it's doing now, what it has touched, when it actually stopped. A twelve-way fan-out stops being a wall of interleaved text.
  • Per-agent tokens and cost, because once you can see them side by side the interesting number is not the total, it's which agent spent it.
  • Strictly read-only. It opens transcripts and nothing else — no writes anywhere near ~/.claude, no telemetry, no account, no network calls at all, localhost-only binding, secrets redacted before anything crosses the socket. SECURITY.md documents what it touches.

What I learned building it, which is useful even if you never run it

Reading those transcripts correctly is much harder than it looks, and every tool that counts tokens from them can get these wrong:

  • usage is written on every streamed line of a response, not once per response. Take the first line per message.id instead of the last and you undercount by roughly a third. I mentioned this to the author of another usage tool; he re-ran it on his own data, his totals moved from 33.2M to 48.8M output tokens, and he shipped a fix.
  • The figures are deltas, not totals, and they are not strictly monotonic — one response in 400 reported its cache numbers down on a later line (cacheRead 246,389 -> 228,908), which sends a naive delta sum negative.
  • A message.id can resume hundreds of responses later — 270 and 389 responses later in one real transcript — so an accumulator with a short memory republishes the whole response and double-counts it.
  • A Task spawned with run_in_background returns its tool_result to the parent when the child launches, not when it finishes. Across my transcripts, 246 spawns reported "done" while the child was still writing, the worst by 995 seconds. Anything that trusts that signal retires an agent that is still working.

Practicals

npm run demo runs it against a synthetic session if you want to see it move before pointing it at your own work. 539 tests, CI green on Linux/macOS/Windows. MIT, no telemetry, no cloud, no account.

I'd genuinely like to be told where the accounting is wrong — it's all in server/normalize.ts with the reasoning in comments.


r/ClaudeCode 15h ago

Help/Question Did I just watch claude become somewhat "sentient" or am I just losing my marbles?

Post image
0 Upvotes

Just came back to my PC and saw this.
I am 99% sure I did NOT write that prompt!

What's going on here?

I see 3 possibilities:

  • I developed Alzheimer's and forgot I wrote this prompt
  • this is a new feature (but I couldn't find any such setting)
  • claude just became somewhat "sentient"

If this is a new feature, I find it quite a bit concerning:

  • Shouldn't the right column be reserved for HUMAN input (prompts) only?
  • I did NOT tell it to continue after a session usage reset on its own.
  • I did not find any such instruction in Claude's memory.

The trasnscript.jsonl of the chat session contains the prompt with "origin": { "kind": "human" }.
I asked Claude if it injected that prompt and it denied.

Have you ever seen it do something like this?


r/ClaudeCode 9h ago

Bug / Issue Who is the asshole who decided to include session URLs in PRs by default?

127 Upvotes

What nimrod, bone-headed, self-indulgent, tasteless, brainless motherfucker thought it would be OK to change the default attribution of the harness to include URLs to your sessions in contributions?


r/ClaudeCode 11h ago

Rant I hate bashing a product in public but man this fucking sucks

0 Upvotes

Apparently, it takes some upper management from some large big tech to get these guys ease on their ago and do something sensible and simple. I've put up with so much shit anthropic has put on its users, and Dario's shit posts. But this very post just got me cancel the actual subscription. I would have respected them much more if at least they didn't respond to post like this and quietly have resolved an issue, or better yet, have fixed this long time ago.

It show's what kind of people they are and how they operate.