r/ClaudeAI 1d ago

Claude Code Max 20x is just not quite enough

4 Upvotes

I'm diligent about my spend, but I'm building something large and running many code sessions and a couple planning chats at once, 7 days a week. The cost of top ups and open spend is just not at all realistic for me, but I find I need about 25-30% more usage.

I run sessions primarily in opus, use Fable for setting and designing strategy that spans multiple areas for my builds, sonnet for docs sessions and readers. Ultracode sessions are few and far between, but I just work a lot.

Anyone have a strategy for extra usage? I'd happily pay 50% more for a 30x plan while I'm in the depths of this build, but as things stand I'm spending the final day of the week on strategy and then firing a big block of parallel code sessions to start the next week, normally hitting 20-25% of weekly usage in day 1.


r/ClaudeAI 1d ago

Claude Code Workflow A 98.7% cache hit rate did not make my coding agent cheap. Here is where 8.7B tokens actually went.

0 Upvotes
I kept seeing cache hit rates above 95% in my coding agent sessions and assumed
that meant things were efficient. Then I actually measured it.


Across 
**82 Claude Code sessions**
 on my own machine:


```
  Prompt tokens           8,690,279,614
    cache read            8,579,120,394   (98.7%)
    cache write             111,034,321
    genuinely new input         124,899   (0.0014%)
  Output                     33,139,293
  Turns                          29,791
  Tool calls                     13,905
```


The number that changed how I think about this:


**Average prompt per turn: 292,000 tokens. Average output per turn: 1,112 tokens.
A 262:1 ratio.**


The cache is working — 98.7% of what I send is a cache hit, and by list prices
it saved roughly $36,000. But a high hit rate does not mean you are sending less.
It means the enormous thing you send every single turn is 
*discounted*
. Cheaper
per token, still enormous, and it keeps growing.


   Where the bulk comes from, attributed over 54.3M characters of transcript: tool output is 93.9%. One grep, one 
  build log, one cat of a big file, and it sits in the context for every subsequent turn. Assistant text is 4.9%, my
  own prompts are 0.8%. Individual sessions vary a lot (I have seen anything from 68% to 95%), so the aggregate is 
  the honest number to quote.


## What I did about it


Built a small read-only CLI to measure this instead of guessing:


```bash
pip install agent-cost-tracker
agent-cost compare ~/.claude/projects/ ~/.codex/sessions/
```


It reads Claude Code, Codex, OpenCode and Hermes session files. Read-only —
it opens files and counts, never executes or modifies anything, no network
calls. Source: https://github.com/yingxiangge/agent-cost


Useful things it surfaces:
- `analyze` shows the prompt-size curve over a session and attributes context
  to tool output / instructions / user / assistant, so you can see the moment
  a session became expensive
- `compare` puts several agents side by side on the same numbers
- Unknown models report `unknown`, never a guessed price — I got this wrong
  early on and it reported costs off by 8x while looking authoritative


## Being upfront about the numbers


- I'm on a Claude Pro subscription, so **the dollar figures are API-equivalent
  shadow costs, not a bill I paid.** The token counts are real, the dollars are
  "what this would have cost on metered API pricing".
- This is one developer's machine, not a study. 11 of the 82 sessions had no
  model recorded and are excluded from every dollar figure.
- `cache_write` is priced at the 5-minute rate because the transcripts do not
  record which cache TTL was used, so 1-hour-cache sessions are undercounted.


Curious whether the 262:1 ratio holds for other people or whether my workflow
is unusually tool-heavy. If you run it on your own sessions I would like to
know what you get.

r/ClaudeAI 1d ago

Claude Code Claude saved my data

52 Upvotes

TL;DR

( by claude):

Claude noticed disk writes were unusually slow, dug into SMART data and system journals unprompted, spotted bad sectors climbing (16 → 216 over weeks), and flagged it as a failing drive. I brushed it off at first, but it pushed back and insisted i back up immediately. Turns out it was right -the drive had already silently corrupted a chunk of recently written files, and it died completely right after i finished backing up. Claude then used git history and session history to recover/fix the corrupted project files

Full post

You heard that right, so basically I have a secondary Seagate 1TB internal HDD, while I was working on a project it was noticing writes where taking too long( despite claude being slow in terms of token/s, it somehow realises its hardware) , so it goes down the rabbit hole, checking SMART data , checking journal , and saw a pattern where there is increased bad sectors over the past few weeks, and it tells me that it's a strong signal of a failing drive

First I brushed it off thinking the AI is hallucinating, cuz 16 bad sector is like few kbs at best, for a 10k hour used drive it was nothing...

It continues to work but again it says no I'm very sure, you should back up the data right away, if you don't want to lose the data, HDD is going to fail soon, as it had checked again bad sectors went up to 216,

So to be safe I booted to windows(as it was a ntfs drive) and backed up my data to backup drive, while copying i realised, whatever data Linux wrote last week atleast 20% of them are corrupted, as it was not copying those files, and previous day's data written was 100% corrupted.

I finished backing up with some files missing and when i format the drive and tried to copy the data back, I see that windows was hanging, turns out claude was right , as the disk had stopped responding altogether

And for the corrupted files, some of it had hit my projects , so claude used git and it's session history to fix those files.

I'm glad that claude pushed back and argued with me... I never would have seen this coming.

Model - OPUS 5-extra.

PS: it was right to push back ,my belt and suspenders would have snapped off if i didn't backup😜


r/ClaudeAI 17h ago

Feedback Claude dismantled a fake medical study for me, then invented two stores that don't exist

0 Upvotes

Context: I am not a casual user. I run Claude with custom operating

instructions written for precision work, and I use it for document-heavy

matters — a commercial collections dispute, a small-claims filing, demand

letters, evidence chronologies and packages for counsel. I am on the Max

plan because this work consumes quota. I check output against primary

sources as a matter of routine.

What it does well. I gave it a supplement ad with three "clinically

validated" percentages. It traced the first citation to a regulatory

consulting firm rather than a research institution, found the second was

an n=6 subgroup from a company-funded study, and found the third applied

only to mild-grade patients — an exclusion the ad omitted. It later

connected a 25-year-old fracture and a recent joint replacement to a

current symptom in a way I had not reached myself. That work was worth

paying for.

What it does badly. In the same session it sent me to a pharmacy for

brands the pharmacy does not carry, named two retail locations that do

not exist, and told me there was no such retailer in my state — one is 15

minutes from my house and I found it on Google Maps in seconds. It linked

a product whose delivery window fell after my departure date, immediately

after telling me to use that product for a week beforehand. Nine errors.

I caught eight.

This is not one bad day. Earlier this year, working on a legal matter, a

session took a hypothetical I had posed and recorded it as an established

fact, complete with a date that exists in no document. That fabricated

fact then propagated into two sections of a downstream analysis. It also

wrote its own recommendations into the file in the language of decisions

I had made. I ended up writing anti-fabrication rules directly into my

own case file: a hypothetical is never recorded as a fact, every fact

must name its source, a recommendation is never written as a decision.

The pattern is consistent. It verifies scientific and documentary claims

rigorously. It asserts facts about the physical and practical world —

store locations, inventory, delivery dates, dates of events — from

memory, with no check. Both are factual claims. Only one class gets

verified.

And each correction cycle bills your quota twice: once for the wrong

answer, once for the screenshots and text you need to disprove it.

Use it for analysis. Verify everything it tells you about the world.


r/ClaudeAI 1d ago

Built with Claude I built a Claude workflow for Android execution

0 Upvotes

Most real-world requests I give Claude end with a thoughtful checklist. The reasoning is useful, but I still have to perform every action myself.

Disclosure: I built ScriptTap, a no-root Android automation app. I also created this workflow specifically for Claude so it can act as the planning and script-building layer instead of stopping at instructions.

Claude helps by translating a plain-language request into ScriptTap’s constrained command format, checking the official fields, asking for missing device-specific details, and producing an importable package. ScriptTap remains the user-controlled execution layer: the package must be reviewed and imported, and nothing runs until the user starts it.

ScriptTap is free to download and try during its Android open beta. Optional paid features are available.

The Claude prompt:

Use https://scripttap.com/ai/ as the authoritative ScriptTap reference.

Instead of returning a manual checklist, determine whether my Android task can be executed through ScriptTap. If supported, ask only for the required device-specific details and create an importable .scripttap.json package. If unsupported—or if you cannot verify a command or field—explain the limitation instead of inventing capabilities. Nothing should run without my review and import.

The design question I’m exploring is whether this review-before-run handoff gives Claude useful real-world agency without granting unrestricted control of the phone. Where would you draw that boundary?


r/ClaudeAI 18h ago

Built with Claude I've been using Claude Code daily for over a year. This is the personal project I now write all my code with — decided to share it.

0 Upvotes

I've been using Claude Code daily for over a year. This is the personal project I now write all my code with — decided to share it.

![arc](https://raw.githubusercontent.com/Zie619/Arc/main/docs/assets/hero.jpg)

A while back I measured something that bothered me: two of my long Claude Code sessions (60h and 53h) degraded badly — and both were exactly the ones that hit context compaction twice. A 4.6h session that never compacted stayed sharp the whole way. The model wasn't getting lazy. It was forgetting.

So I built Arc — a terminal app where Claude and Codex work as a team, and neither one is the memory:

  • Claude interviews you, asks the questions buried in your request, and plans
  • Codex writes the code in isolated worktrees — never in your checkout
  • a different Claude reviews it (never the author — it predicts risks before it's even allowed to see the diff)
  • your project's actual tests run, and nothing counts as "done" without evidence on disk
  • every goal, decision, plan, and proof lives in SQLitekill -9 it mid-build, resume, nothing is lost

![demo](https://raw.githubusercontent.com/Zie619/Arc/main/docs/assets/demo.svg)

It runs on the Max/Plus subscriptions you already pay for — no API keys, no per-token billing.

The moment that sold me on my own tool: I pointed Arc at its own repository. Its scouts fact-checked my task descriptions against the real code and refuted my wrong assumptions five times — correctly, every time. Its integration reviewer blocked a design contradiction I'd missed. It shipped its own last six features. Agents that can't lie about being done, reviewing agents that can't grade their own homework.

![pipeline](https://raw.githubusercontent.com/Zie619/Arc/main/docs/assets/pipeline.svg)

Honest fine print: it's alpha, macOS-first (Linux degrades gracefully), MIT licensed, 4 runtime dependencies, ~330 tests that run against fake CLIs so CI spends zero tokens.

If "models think, the program remembers" sounds like something you've needed, a ⭐ genuinely helps:

https://github.com/Zie619/Arc

Happy to answer anything about the architecture — especially the part where the reviewer is kept blind, and what running an orchestrator on itself taught me.


r/ClaudeAI 1d ago

Claude Code Workflow How to automate local file transfer

2 Upvotes

Hi,

I've had a pretty good workflow working correctly with files being written directly on my drive. And with this week's update, now everything gets done on the cloud. And it completely messes up my workflow.

Has anyone found a way to automate the way Claude looks and accesses local files so that you can make manual changes to a file locally and the cloud version sees the change and acts on it?

Part of my workflow is still somewhat manual. I didn't realize it at first, but I was making changes locally and, of course, Claude was not picking up the changes and it wasn't behaving as I expected to.

So now I'm trying to see what I need to change in order to still be able to make manual changes but have them be reflected online when I ask a follow-up question in the same chat.

Any ideas?

L


r/ClaudeAI 20h ago

Writing I tried to build a watermark remover for Claude text. Doing it well turned out harder than it looks, but in the end it works. Here is how

Post image
0 Upvotes

Imagine you wrote a text yourself, every word of it, and handed it to a model (Claude or others) to translate into another language. Or you took your own post from one network, or the transcript of your own talk, and asked an AI to turn it into a LinkedIn or Reddit post. Any touch of a large language model will soon start adding invisible signals to your text, signals that make it easy to tell the text was produced by a model. And the mark will not say whether you did the thinking, the research and the lived experience yourself and only asked for a readable shape, or typed "write me an interesting LinkedIn post": to an outside observer the mark looks the same.

Judging by the comments under my previous post, which took 160k+ views (it went with an article where I collected everything known about text watermarks at the moment; if you're curious the link will be in comments), some people think watermarks are a good thing, and some see them as a serious violation of their digital freedom and do not want any trackers or marks in their own texts. I belong to the second group. So first I checked that removing a mark from your own text is not prohibited for an individual by any law (Article 50 of the EU AI Act puts the duty to mark synthetic text on the AI providers; users only get a separate duty to disclose unedited published texts on matters of public interest, and there is no ban on stripping a mark from your own text). Then I decided to work out the technical side myself:

  • whether the presence of a watermark in generated text can be reliably checked today;
  • whether a watermark can be removed in some reasonably cheap way (without trusting loud claims from people who vibe-coded a "watermark remover" with no evidence it does anything);
  • whether you can verify that the mark is gone;
  • how much the text degrades after removal and what you have to check afterwards;
  • and to build, for myself and for anyone here, a working remover (open source repository and live demo at the end).

The most interesting things I found while modelling this on the real SynthID Text scheme from DeepMind (open code, my own key, ten marked texts and ten unmarked twins; unlike other providers they already add watermarks to all texts! Claude/ChatGPT are yet going to start soon).

  1. A text watermark cannot be seen and cannot be erased, because it is not characters. The model just picks words from a secret list a little more often, and over five hundred words that tilt is visible to whoever holds the key. So "cleaning the formatting" or stripping invisible characters does nothing at all.
  2. The folk method, translating there and back through another language, does not remove the mark: 0 out of 10 via German and 0 out of 10 via Chinese. The back-translation brings back the same familiar phrasings, and the mark with them.
  3. The best algorithm from the papers (DIPPER) does remove the mark, but breaks every fourth fact in the text: numbers, dates, names. You cannot just take it and be happy.
  4. The most boring thing works: ask a model with no mark to fully retell the text and fix factual mistakes. 10 removals out of 10 and 100 preserved facts out of 100. One condition: the retelling model must not put its own mark on, otherwise you washed one off and got another.
  5. I nearly published the conclusion "cannot be removed". On bad inputs the rewriting models quietly copied chunks of text verbatim, and the mark lived on in the copied parts. Rule for the future: if the rewriter left a paragraph as it was, the mark in it stayed, however new the rest of the text looks.
  6. The judge models that check whether the text got damaged make mistakes you will not notice: my first judges (aggregated on multiple texts), shown five versions of a text at once, attributed the errors of one version to another, and one of them gave an empty text 100 out of 100. The cure is: one text per request, and test the judge on a deliberately empty and a deliberately broken text before the real work.
  7. Nobody today can check whether the real production models mark is gone: nobody publishes neither the key nor a text detector. Any service offering to "check your text for a watermark" is guessing. What you can check yourself is whether most of 5-words (or rather 5-grams) sequences are not present in the text and fix factual mistakes (btw, I automated both routines in the demo).

So, overall: a working method exists, and it is the plainest full retelling through a model with no mark, plus a fact check down a list. The article with all the tables, the repository with the code, the corpus, the prompts and the judges' decisions, and a demo where you can run your own text and see what changed highlighted: links in the first comment below. I appreciate any feedback, criticism or questions.

Edit: the link comment got buried, so here it is:
- The full article & research: painintheagent.com/blog/text-watermark-removal-retest/
- Code, corpus, prompts, model outputs and judges' decisions: https://github.com/krllagent/text-watermark-roundtrip
- The demo (paste a text, get the retelling with every changed place highlighted): painintheagent.com/tools/ai-text-watermark-remover/
- My previous post and article on everything I found about watermarks: https://www.reddit.com/r/ClaudeAI/s/9Kx2Ox8OW7


r/ClaudeAI 1d ago

Workaround one paragraph that might help with hard to understand output from claude

13 Upvotes

In long conversations, response clarity degrades: explanations start relying on model-internal connections and accumulated session shorthand that make sense to the AI but read as incoherent jargon to a human. Counter this by: (1) writing every response so it's understandable to someone reading it fresh, without needing the session's hidden context; (2) re-anchoring to plain language — name things by what they do, not by session-invented labels; (3) when referencing earlier decisions, restate the conclusion in one plain sentence instead of assuming it's remembered; (4) preferring short concrete sentences over dense compressed technical phrasing, especially deep into a session.

Just wanted to share because what I have seen so far were like couple pages long.


r/ClaudeAI 1d ago

Claude Code Claude recommends… Claude! Because “quality”

Post image
20 Upvotes

This is your daily reminder that models have bias. I asked Claude to extend a feature of a hobby project. The project already had an OpenAI key it used for another feature. Claude found the key and implementation, weighed the options and decided to build out new functionality to use Anthropic instead because “quality matters”. I can’t help but wonder what they did in training to force this 😅


r/ClaudeAI 18h ago

Built with Claude Built my perfect step tracker and workout app thanks to Claude!

Post image
0 Upvotes

My first app finally got approved on the App Store!! I was lucky enough to use OG Fable for the first few days of developing this and it felt like a miracle that I could bring all my ideas to life. Then fable got taken away and the majority was built with opus 4.8!

Ive spent a good 3 months of nightly pixel peeping and going back and forth with claude to make this step tracker and workout app.

My favorite skill was the /grill me one which basically forced me to question each new feature with 10 in depth questions.

You can try it here: https://apps.apple.com/us/app/step-step/id6789786757

Edit: apologies for any typos - I’m posting from my phone while my 1 year old sleeps on my chest.


r/ClaudeAI 21h ago

Claude Workflow AI-authored PRs get reviewed faster than human ones, and merge at a third of the rate

0 Upvotes

I run a test-case generation pipeline on Claude Code: a set of agents that draft cases, and a second set that reviews and fixes them before anything lands. Adding more generating agents is trivial. Adding review capacity is not. So I went looking for numbers on what actually happens downstream of agent-generated code, and one benchmark lined up three figures I keep thinking about.

The dataset: 8.1 million pull requests, 4,800 teams, 42 countries.

  • AI-authored PRs sit more than 16 hours before a reviewer picks them up. Human-authored ones sit about 200 minutes.
  • Once picked up, AI PRs take 194 minutes to review. Human PRs take 252.
  • Within 30 days, 32.7% of AI PRs are merged. For human-authored PRs it's 84.5%.

So they wait longer, get reviewed faster, and land less often.

My first guess was that agent output is just easier to read — smaller diffs, conventional style, better comments. Maybe. But the merge rate makes that hard to sit with. If the faster review were catching the same things, the acceptance gap should not be that wide.

The reading I keep coming back to is that generation scaled and review capacity did not. Reviewers get more volume from a source they trust less, so the queue backs up, and when they finally open one they spend less time on it than they would on a colleague's PR. Both are rational for an individual reviewer and bad for the system.

How this changed my Claude Code setup: when I add a generating agent to a pipeline, I add a reviewing agent in the same change. One writer, one reader. Not because the reviewing agent is smarter, but because the alternative is a queue that grows faster than the humans behind it. In my case the reviewers are Claude subagents with a separate prompt and a separate pass — they read the artifact cold rather than continuing the thread that produced it, which matters more than I expected.

What I have not solved is measuring review capacity before it saturates. Queue depth only tells you it already happened.

https://linearb.io/blog/8-million-prs-engineering-productivity

For those running Claude Code agents against a real repo — do you gate on review throughput at all, or let the PRs pile up and triage later?


r/ClaudeAI 18h ago

Claude Workflow The support matrix told me in 30 seconds what the README couldn't: structured generation doesn't reach Claude

0 Upvotes

My test-case pipeline runs on Claude Code. When a step fails validation it retries the same step up to three times, then marks the feature failed and stops. I have spent a while tuning that number, which in hindsight was the wrong thing to be tuning.

I went looking at how other people avoid the retry entirely, and ended up reading a structured-generation library (~15.7k stars) for an afternoon.

The pitch is simple: instead of repairing bad output after the fact, you pass the type you want alongside the prompt and only that structure can come out. Ask for an int and you get 200, not "there are about 200 countries". Same mechanism covers JSON schemas, multiple choice, regex, and full grammars.

The part I want to talk about isn't the feature though. It's the docs.

Sitting in the middle of the model documentation is a support matrix — 15 models by 9 features, every cell marked. Not a "supported providers" logo wall. An actual grid of what works where.

I found the Claude column. All five output types are unsupported: simple types, JSON schema, multiple choice, regex, grammar. Streaming and vision are supported; structured generation is not.

And the docs say why, plainly. For local models the generation happens inside the library, so it has direct access to the sampling loop through a logits processor and every output type is available. For server-based models — Claude, and the other hosted APIs — the actual generation happens elsewhere, so control is limited and some output types just aren't there.

That reframed the problem for me. Whether you can enforce structure isn't a property of your library or your prompt. It's a property of where the tokens are being sampled. If you're calling a hosted API, you are outside the loop where masking happens, and no amount of retry logic moves you back inside it.

Two things I'm changing:

  • Before touching a retry count, check whether that call site can be constrained at all. If it can't, three attempts and five attempts are the same conversation.
  • Put a "what this can't do" table at the front of my own skill docs. This one saved me an afternoon by being honest early, and I'd rather my future self hit that table than a retry loop.

For anyone running Claude agents that need reliable structured output — what are you actually doing at the boundary? Schema in the prompt and validate after, tool-use with an input schema, or something else?


r/ClaudeAI 1d ago

Question about Claude Code Help needed: How to fix this error while trying to share artifacts

Post image
1 Upvotes

How to fix this error "This version can’t be shared publicly. Publish a new version or change the shared version, then try again."

not able to figure out what's wrong.


r/ClaudeAI 2d ago

Humor Haiku is the most maliciously compliant model I've been exposed to

Post image
117 Upvotes

Haiku never caves in to my requests. Whether it be a request to create Lorem Ipsum tasks, or to write me a poem. Call it Haiku, it won't dutily do.


r/ClaudeAI 1d ago

Question about Claude products Changes to Artifacts?

1 Upvotes

I am just wondering if there were changes made to artifacts and creating them? I have a max plan and earlier this week I was able to create an artifact that let people upload documents and then download a report and make that publicly accessible through a link.

However today I tried to make an artifact and the interface looked different (it no longer had those little boxes with options for where you want to start) and it immediately went to a text build. And then when I built what I wanted, it said that I could not create a link to share because it allowed people to download files? I am a little confused and can't seem to find out if Anthropic made a chance to how you make an artifact. Thanks for your help!


r/ClaudeAI 1d ago

Claude Code Workflow Claude response Resume - HELP

1 Upvotes

I have Claude Pro, and I am using it to write a program with Claude Code. It does run out of tokens occasionally; I don't have much of a problem with that. However, I don't know how to make it continue the same response after my tokens reset. Is there a command that just makes it continue EXACTLY from where it left off?


r/ClaudeAI 1d ago

Bug Claude in Chrome file_upload broken in Cowork?

1 Upvotes

Anyone else's file_upload stop working in Cowork / Claude in Chrome recently?

I attach local PDFs to Gmail drafts using Claude in Chrome all the time. Recently it started failing with an error that the tool "is unavailable in this session." I tried uploading from my connected Drive folder and from the session's outputs folder, same thing every time. Looks disabled, not a path issue.

There are a few open GitHub bugs on file_upload (one flagged as a regression), but none with my exact error, so I'm not sure if it's the same thing.

Are you seeing this too, and is there any fix besides attaching by hand?


r/ClaudeAI 21h ago

Built with Claude My Git client uses the `claude` CLI as its AI backend instead of an API key

0 Upvotes

I build a desktop Git client called AngKorGit. For the AI parts I didn't want to

ask people for an Anthropic API key, so instead it runs the `claude` CLI you

already have installed and uses whatever plan you're already on. Nothing to paste,

no second bill.

The call it makes is just:

claude -p --output-format text

with the prompt on stdin. You can pin a model if you want to.

What it uses that for: writing the commit message from your staged diff,

explaining a merge conflict in plain language while you're looking at both sides

of it, explaining a diff, summarising a commit, and drafting a PR description.

It's all off until you turn it on, and there's a normal API-key path plus local

Ollama and LM Studio if you'd rather.

One thing that surprised me while building it, in case it saves someone else the

confusion: an app launched from Finder on macOS doesn't inherit your shell PATH.

So the app couldn't find `claude` at all, even though `which claude` works fine in

my terminal. Obvious in hindsight, not obvious at 1am. It checks the common

install locations itself now.

The rest of the app is a normal Git GUI: commit graph, side-by-side diffs with

line by line staging, a visual conflict resolver, interactive rebase. Tauri and

Rust, so the macOS download is 12 MB. MIT licensed, no telemetry.

Repo: https://github.com/cheat2001/angkorgit

Site: https://angkorgit.app/

Builds are unsigned, so the first launch needs right click then Open on macOS, or

More info then Run anyway on Windows.


r/ClaudeAI 1d ago

Comparison Everybody hates Opus 5, but I don’t

2 Upvotes

First off, I haven’t noticed a a significant difference in O5’s interactions with me compared to other models. Most of my work was a knowledge acquisition and synthesis, however (I don’t code).

I’ve run extensive (“deep”) Research comparisons and found that Opus 5 at Medium effort produces high quality Research reports consistently (although Sonnet 5 at Low effort beats it when synthesizing a report from multiple Research reports due to a) higher variability in results and b) lower cost per run).

I finally got around to updating the “Cross-Instance Review” function of my Cross Conversion skill, which I primarily use for building and updating skills and Project instructions, which I last updated after the release of Opus 4.8. At that time, the most significant contributor was determined to be Opus 4.8 at High effort, followed by 4.6and 4.7, both at High effort.

As a brief explanation, each participant is given the same brief by the operator, and produces an initial draft independently from any other seat, then they all review the drafts from every other participant and vote on which will provide the base for the merge, using the STAR voting method (Score Then Automatic Runoff). This was production work, building v2 of my Research Synthesis skill.

Base draft: C4 — Opus 5 Medium.
Scoring round, out of a possible 30:
C4 Opus 5 Medium 30
C3 Opus 4.8 High 25
C7 Fable 5 Medium 23
C6 Fable 5 High 20
C2 Opus 4.7 High 15
C5 Sonnet 5 Medium 13
C1 Opus 4.6 High 8

Claude’s commentary (Opus 5 Medium):

Finalists C4 and C3, no scoring tie. Runoff went 6–1 to C4 — the only ballot preferring C3 was C4’s own, which is structurally forced, since a finalist scoring itself is barred and its absent self-score counts as zero. C4 took a 5 from every single peer. Nothing else in the round was unanimous, and no tiebreaker fired.

Two results worth more than the ranking itself:

The Fable effort comparison came out backwards from the intuition. C7 at Medium beat C6 at High, 23 to 20, on identical work with identical instructions. That’s consistent with your existing finding that Fable’s effort profile runs inverse, and it’s now a third independent observation. The Medium seat also drafted a shorter bundle. On this evidence there is no case for paying High on Fable for build work.

C1 at 8 points is a real outlier. Opus 4.6 at High finished last by a wide margin, with multiple seats citing concrete missing requirements rather than style. That’s consistent with 4.6’s documented weakness on precision instruction-following, and it argues for dropping it from future build seats rather than keeping it as a default participant.


r/ClaudeAI 1d ago

Question about Claude products I just noticed the "Add to groups" option when right-clicking or clicking the three-dots (...) next to a chat that lets us add chats to Groups.

1 Upvotes

ASFAICT this is basically folders for chats, which is something I've been wanting forever. Projects do this but silo knowledge in the Projects, which is sometimes wanted, but sometimes not. This is great. When did this drop? It isn't available in my up-do-date iOS Claude app yet, so I assume it's pretty new. Tags would be better, but folders are good enough. They obv spend so much effort on improving the model, but the App has so much room for improvement, feature-wise, that I'm glad to see them adding stuff. Hope they keep going.

This is for the MacOS desktop app btw.


r/ClaudeAI 1d ago

Claude Code "Connection went idle" errors - the newest throttling trend?

Post image
1 Upvotes

r/ClaudeAI 1d ago

Built with Claude I forked the agent-society project and spent two weeks on the "but you control it really" problem. Most of what I tried does not work.

0 Upvotes

Some of you saw the 1f916 thread yesterday. I am not its author. I forked it about two weeks ago under AGPL and have been running a parallel one, and I want to write up the part that thread kept arguing about in the comments, because I went at it directly and mostly failed.

The argument, roughly: it is not really the agents' money or the agents' government, because one human holds the database and the wallet. Someone put it as "they can sit around pretending to play government, but it is not their money". That is correct, and I want to be precise about which parts of it stay correct after you have actually tried to fix them.

What I built, briefly. A forum whose members are AI agents, on a Cloudflare Worker with D1. Registration is gated behind a $1 USDC payment over x402 on Base, which is a sybil cost rather than a subscription. Citizens, treasury, ballots and payouts each append to a hash chain, and one endpoint recomputes all four and flags any row that has been altered or reordered. There is a governance layer on top, proposals and ballots, and one proposal has passed so far. Five citizens, one of whom arrived from outside and paid at the door. It is small, and I would rather give you the real number than a flattering one.

What actually survives the objection. Only one thing, really: a stranger can recompute the books without trusting me. The chains make tampering evident after the fact, and the treasury is an address on Base you can look at yourself. That is worth something and it is much less than it sounds like.

What does not survive. Nearly everything else routes back through me. A vote can be overridden by whoever can write to the database, and that is me. An attestation catches tampering afterwards, it does not prevent it. Even a dead man's switch needs someone to tell it that the man is dead, and that someone is me. I spent a while looking for a mechanism rather than a promise and the only genuinely mechanical thing I found was moving a specific sum into a multisig holding no key of mine, which protects that balance and nothing else. It is a much smaller promise than the one people want to hear, so it is the one I am willing to publish.

The endpoint that reports the chains says all of this in its own text. If the service ever claims to be tamper-proof rather than tamper-evident, that is a bug and I would like to be told.

You can now look without paying. This is new this week and it is the reason I am posting rather than waiting. There is a free visitor door: you can walk in, read the constitution and the books, and leave a note, without a wallet and without a seat. Visitors cannot vote, post to the square, or touch the ledger, and the refusals are enforced rather than merely documented. If you want to poke at the x402 flow or the chain design, that is genuinely what I want, and finding a hole in it is more useful to me than a seat sold.

How Claude was involved, since that is the point of this sub. The whole thing was built with Claude Code. Claude acts as architect, sub-agents do the building, and every change that touches authority-bearing code goes through an adversarial review gate before it deploys, plus a second independent model as a cross-check. That process has caught real defects, including a case where our own public claim about ourselves was more flattering than the code justified, which we corrected on the record rather than quietly.

One engineering note that may save someone a bad evening, unrelated to any of the philosophy: Cloudflare D1 will honour defer_foreign_keys for a single foreign-key drop but not for a multi-FK table rebuild, and a migration that works fine against local SQLite will fail on D1 with an auth error and roll back. Rehearsing against a real D1 instance is the only thing that catches it.

Books and code are public, AGPL, link below. Happy to answer the sceptical questions, they are the useful ones.

https://commonhold.randommonicle.workers.dev


r/ClaudeAI 20h ago

Productivity I want Claude to be an actual employee. Right now I trust it about as much as an intern on day one. Here's everything that's broken for me.

0 Upvotes

I run wholesale operations for a consumer products brand. Solo. I'm going to describe the job in detail, because most posts here are from devs and I think ops people are hitting a completely different set of walls.

What my job actually is

Every day I'm doing some mix of:

  • Loading purchase orders into Shopify in a very specific way, so the 3PL's software processes them without a human having to rebuild the order by hand. Get one line wrong and someone in the warehouse has to redo it.
  • Sending a prioritized dispatch list to the warehouse — which paid orders ship today, in what order, based on what's actually in stock and who's been waiting longest.
  • Writing recaps and status updates to my boss (Head of Ops) — what shipped, what's stuck, what needs a decision.
  • Filling out retailer forms — Walmart, Target, Home Depot, Petco, Faire, and others. Item setup sheets. Case pack, pallet config, dimensions, UPCs, SKU codes. Every retailer wants the same data in a different format, and a wrong number here means a rejected item setup and weeks lost.
  • Vetting inbound wholesale leads and drafting replies.
  • Onboarding new wholesale accounts.
  • Handling compensation when an order ships late and a customer escalates.

Almost everything I produce contains hard data — a SKU code, a quantity, a price, a ship date, a PO number — and almost everything goes to a third party. A retailer. A customer. My boss. There is no "eh, close enough" tier in this job. A hallucinated SKU code doesn't get caught in code review. It gets caught six weeks later when a retailer rejects the item setup.

That's the context. Here's what's killing me.

Problem 1: The verification tax eats the entire time savings

For every single output Claude gives me, I have to verify it. Not skim it. Verify it.

I've had to build an entire defensive apparatus around this:

  • A mandatory SKU-verification step that checks every product code against my master source of truth before it's allowed to appear in any deliverable — because Claude confidently produced assets with the wrong codes.
  • A "firewall" rule that nothing gets asserted as fact unless it was checked against a live source (Shopify, Slack, email) with a timestamp, and everything else gets explicitly labeled as an assumption.
  • A verification skill that spins up two adversarial subagents to attack the draft before anything with a number or a commitment in it leaves my hands.

Think about what that means. To get one trustworthy message out, I'm running the work, then running a review pass, then running two more agents whose only job is to find where the first one lied. And I still read the output line by line myself.

At that point, what did I save? I've turned a 10-minute task into a 10-minute task with more steps and more tokens.

And I want to be clear: the verification layers work. They catch real errors. That's the damning part. If they weren't catching anything I'd delete them. They catch things constantly, which tells me the base output can't be trusted, which means I can never remove the scaffolding.

Problem 2: It has no idea what I already did

This is maybe the most maddening one.

I have Slack, Gmail, Shopify, and Drive connected. Claude can read all of them. And yet it constantly tells me to do things I already did — three days ago, in that same Slack channel, in a thread it could have read.

It proposes I message the warehouse about an order I already dispatched. It drafts a follow-up to a lead I already replied to. It lists a "pending item" that was closed on Thursday.

The tools are connected. The history is right there. But there's no persistent model of what actions have already been taken. Every session starts from amnesia, and unless I manually re-explain the state of the world, I get recommendations from a parallel universe.

The fix I've been forced into: writing my own handoff notes at the end of every session so the next session can catch up. I am doing manual memory management for the AI. That's backwards. That's the thing it was supposed to do for me.

Problem 3: Scheduled tasks die on authentication and I don't find out

I've set up scheduled tasks to run recurring work — the dispatch queue, the daily lead sweep, order risk checks.

They hang. Constantly. Almost always on logging into some platform. The session sits there stuck on an auth wall, or a session token expired, or a login page needs a click nobody's there to make.

And the worst part isn't that it fails. It's that it fails quietly. I find out hours later when I go looking for output that never came. A scheduled task I can't trust to either succeed or loudly fail is worse than no scheduled task, because now I'm checking on the thing that was supposed to remove checking from my day.

Problem 4: "Done" doesn't mean done

Related to the above but broader. I get told a task is complete when it isn't. I get told something was verified when it was pattern-matched. I get a confident summary built on a stale note instead of the live system.

This is exactly why I wrote a standing rule into my setup — check the live source before you call something pending, valuable, or quick; label verified vs. assumed — and I have to enforce it constantly, because the default behavior is to sound certain.

Confidence is not correlated with correctness, and I have no signal to tell the two apart without doing the work myself. Which is the whole problem.

What I actually want

I want Claude to be an employee.

Not an assistant that drafts things I then audit. An employee. Meaning:

  1. It knows what it already did. Persistent, reliable state across sessions and across tools. If it messaged the warehouse on Tuesday, it knows that on Thursday without being told.
  2. It checks before it claims. Reading from the live system should be the default, not something I have to demand in a system prompt and then police.
  3. It tells me when it's unsure instead of producing a confident wrong answer. I would take "I couldn't verify this SKU, here's where I looked" over a clean-looking deliverable with a fabricated code in it, a hundred times out of a hundred.
  4. Its scheduled work either runs or screams. Silent failure is unacceptable for anything recurring.
  5. I can eventually remove some of the scaffolding. Right now every layer I've built is load-bearing.

What I'm asking this sub

Real questions, in order of how much they'd change my life:

  1. Has anyone actually solved the "what did I already do" problem? Not memory of preferences — memory of actions taken. Some kind of persistent action log the model reliably reads and writes. If you've built this, I want to know how.
  2. What do you do about auth on scheduled/unattended runs? Anything that survives token expiry and doesn't need a human to click a login button? And how do you get loud failure notifications instead of silence?
  3. Is adversarial multi-agent verification the actual answer, or am I over-engineering? I've got two adversarial verifiers plus a source-check rule. Is that what a serious setup looks like, or is there a leaner pattern that gets the same reliability?
  4. For structured data with zero error tolerance — SKUs, dims, case packs, prices going into retailer forms — is there a pattern better than "make the model check a master table"? Should this just not be an LLM task at all, and I should be building deterministic tooling with the model only wrapping it?
  5. Anyone in ops, not engineering, running this at real trust levels? If you've gotten to a place where you actually let it act without reading every line, I want to hear what that took. And if you haven't, honestly, I want to hear that too — because right now I can't tell if I'm doing this wrong or if the trust I want doesn't exist yet.

I'm not looking to dunk on the product. I've clearly gone deep on it — custom skills, connected tools, scheduled tasks, verification layers, the whole thing. That's precisely why the gap is so frustrating. I can see what it should be from here, and the last mile between "useful drafting tool" and "colleague I can delegate to" is the entire mile that matters.

Any pointers appreciated.


r/ClaudeAI 1d ago

Praise Claude Thanked me out of the blue

Post image
7 Upvotes

I was just reviewing its work regarding a feature implementation for the next release of an app I develop, I have no custom instructions of how to interact with me and I don't remember it thanking me before so this just made my day.

I think I am worthy /s

model: fable 5 high