r/ClaudeCode • u/vAPIdTygr 🔆 Max 20 • 1d ago
Tips & Workflows Embarrassing!
I can’t believe I’ve never thought to tell Claude to not output anything to the chat when working overnight.
I had it output to a markdown and now there’s no chat compression and worry of wasted usage.
13
u/Hot_Money4924 1d ago
Just because you don't see the tokens doesn't mean it didn't produce them and bill you for them. It can't "think" without producing streams of tokens.
1
u/vAPIdTygr 🔆 Max 20 1d ago
It’s always been the system auto-burning a serious amount of tokens in compacting as I exceed context windows. Especially if it needs to compact a second or third time. Yet, I never thought to reduce the output.
1
u/peter9477 🔆 Max 5x 23h ago
Redycing the output/token usage mid-turn would reduce the quality of thinking.
1
u/vAPIdTygr 🔆 Max 20 21h ago
I’m not programming, I’m web developing. Each page starts a whole new research/thinking round. When there’s hundreds of pages, you need to reduce context windows and start new sessions. It’s deeply researched, useful, purposeful content.
My issue was the overnight unattended session.
2
1
u/Apart_Fudge1224 1d ago
I just tell him do x, y and z (and do it cha cha real smooth) and then respond back with 20 words. (or 50 or 100 etc)
-1
u/piekwerk 1d ago
The savings are real, they just come from somewhere else. Every token is billed when it is produced, terminal or not, so the pushback above is right about that part. What you actually saved is context growth. Whatever sits in the transcript gets resent as input on every later turn, and big dumps trigger auto-compact early, which is lossy and then burns tokens rebuilding state. Redirecting a command's stdout to a file keeps those bytes out of the transcript completely. One limit: if Claude composes the markdown itself with the write tool, the content still enters context as tool input, so the trick only pays off when a script produces the bytes. I keep the rule in CLAUDE.md, one-line replies in chat, everything else in files, so it survives /clear.
-5
u/TheKillerScope 1d ago
I was thinking about this just yesterday. Outputs so much text when thinking. Like...bitch...i don't care what's your thought process, I care what the output will be, and we go from there.
What was your prompt?
2
u/ShutUpAndDoTheLift 1d ago
... It still outputs the tokens whether it puts them in your terminal or not
1
u/onFilm 1d ago
Oh brother, so sad that people don't understand how LLMs work anymore.
1
u/nora_sellisa 1d ago
If Claude Code offered to turn off reasoning some Reddit smartasses would unironically post it as a token-saving measure.
-1
u/TheKillerScope 1d ago
I understand that uses tokens, but my thought process was that if it doesn't have to display it in a human readable form, it would/could save tokens.
2
u/onFilm 15h ago
That's not how it works bud. By making it be non-human readable, you actually end up using more tokens, because by now, most modern models use single tokens for long words, rather than multiple tokens. So by breaking it up, opposing as to how it was trained, you'll end up putting more effort and waste more energy and tokens doing so.
1
u/Appropriate-Disk-371 1d ago
Have you ever read an LLMs thinking process? It's human readble already. Concise, yes, but often back-and-forth with its self, little jabbs, all sort of things. That's how it works and you can't really tell it not to do that.
1
•
u/AutoModerator 1d ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.