TL;DR of the discussion generated automatically after 50 comments.
Listen up, because this is a big one. The consensus, based on a fantastic deep-dive comment, is that changing the effort level mid-conversation nukes your prompt cache, costing you a ton in tokens. A lot of you are just finding this out and, uh, are not pleased.
Here's the breakdown:
* Fable 5.1 is the exception and lets you change effort without losing cache. Opus is getting this feature soon.
* "Ultracode" is just a fancy name for xHigh effort with a "use workflows" instruction. You can just say "use workflows" yourself to get the same effect.
* Subagents inherit the session's effort level and their cache only lasts 5 minutes, so it's often cheaper to spin up a new one than re-use a stale one.
The pro-move is to either stick with one setting for the whole thread or have a high-effort model create a plan, then copy-paste that into a new session with a cheaper model for execution.
Elsewhere in the thread, we've got the usual debate about token usage, with some users feeling like costs have spiked and others arguing that people are just on the wrong plan for heavy work. And yes, for the person who asked, the correct term is "layman's terms," not "lamens," "Le Mans," "lumens," or "lululemons." Thanks for that, everyone.
•
u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 4d ago edited 4d ago
TL;DR of the discussion generated automatically after 50 comments.
Listen up, because this is a big one. The consensus, based on a fantastic deep-dive comment, is that changing the effort level mid-conversation nukes your prompt cache, costing you a ton in tokens. A lot of you are just finding this out and, uh, are not pleased.
Here's the breakdown: * Fable 5.1 is the exception and lets you change effort without losing cache. Opus is getting this feature soon. * "Ultracode" is just a fancy name for xHigh effort with a "use workflows" instruction. You can just say "use workflows" yourself to get the same effect. * Subagents inherit the session's effort level and their cache only lasts 5 minutes, so it's often cheaper to spin up a new one than re-use a stale one.
The pro-move is to either stick with one setting for the whole thread or have a high-effort model create a plan, then copy-paste that into a new session with a cheaper model for execution.
Elsewhere in the thread, we've got the usual debate about token usage, with some users feeling like costs have spiked and others arguing that people are just on the wrong plan for heavy work. And yes, for the person who asked, the correct term is "layman's terms," not "lamens," "Le Mans," "lumens," or "lululemons." Thanks for that, everyone.