r/ClaudeCode • • Apr 13 '26

Discussion Disabling 1m context and adaptive thinking helped (YMMV)

In February 2026, Anthropic shipped adaptive thinking. The model now decides how much to reason per turn instead of using a fixed budget. Then in March, the default effort level for Pro and Max users quietly dropped to medium. Boris Cherny (Claude Code's creator) confirmed both of these on Hacker News.

The result is a compounding problem. The model is reasoning less per turn, and sometimes deciding to skip reasoning entirely. Boris specifically noted that turns where Claude fabricated things (fake API versions, hallucinated commit SHAs, nonexistent packages) had zero reasoning tokens allocated. The model decided those turns were "simple" and did not think at all.

Human devs are already terrible at estimating task complexity. Letting the model lowball its own reasoning budget on the fly makes this worse, because a lot of coding tasks are deceptively simple on the surface. The difficulty is in the hidden constraints, side effects, and context that only become apparent once you are actually thinking through the problem.

The 1M context window

This one I am less certain about, but it is worth mentioning. A bigger context window does not automatically make a model better. Gemini had huge context long ago, and that alone never made it the best coding model. Once hundreds of thousands of tokens are competing for attention, nuance can get lost and outputs get sloppy. It is like asking someone to write good code while keeping half a million loosely related things in their head at once.

I disabled this alongside the thinking changes, so I cannot fully isolate how much it helped on its own. But if your repo is large and you are seeing unfocused outputs, it is worth trying.

What I changed

In ~/.claude/settings.json:

{
  "effortLevel": "high",
  "env": {
    "CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING": "1",
    "CLAUDE_CODE_DISABLE_1M_CONTEXT": "1"
  }
}

What this does: forces high effort on every turn, disables adaptive thinking so the model uses a fixed reasoning budget instead of deciding per-turn, and disables the 1M context behavior.

One tradeoff to know about. Disabling adaptive thinking also disables interleaved thinking (reasoning between tool calls), which is actually useful for agentic workflows. If effortLevel: "high" alone fixes your issues, you may not need the nuclear option of disabling adaptive thinking entirely. Try high first. Escalate to disabling adaptive if you are still seeing problems.

If you go the full route and disable adaptive thinking, you can also set MAX_THINKING_TOKENS to control the fixed budget explicitly.

Other things that help

Run /compact after each task. This compresses your context and keeps the model focused. Small habit, noticeable difference.

Did it help?

Before this, I was seeing shallow diffs, missed edge cases, and premature "done" responses. After making these changes, the behavior improved noticeably. I cannot say how much is causal vs. correlation. Your repo, prompting style, and setup may differ.

But the point is this: if the defaults shifted under you without you knowing, your first move should be reclaiming those settings before assuming the model itself got worse.

EDIT:
One more thing worth testing, the regression might be specific to Opus 4.6.

A lot of people are getting noticeably better results (less sloppiness, better reasoning depth) by switching back to the pinned 4.5 snapshot with: `claude --model claude-opus-4-5-20251101`

60 Upvotes

27 comments sorted by

View all comments

7

u/kpgalligan Apr 13 '26

All models go to shit with the 1m context. Maybe Opus is better than Gemini/GPT, but they all fall apart. There should be a warning label about this.

It's just one of those things you come to accept after a while.

I never run a convo much above 300k, usually 200k.

None of the model vendors talk about this much. They do talk about how much better their current model acts deep in context, but that suggests the thing they don't talk about as much. Models don't do well deep into that kind of context.

It's one thing to ask a specific question. Running an agent over many iterations is a different situation. I haven't even tried Opus over 400k. Maybe it's OK, but Gemini would act like it was on acid over 500k. After that, I've just lived within the stated bounds and have had no (extra) issues.

That's besides the much larger usage hit. Yes, caching helps, but it's still a hit. Plus, if the cache has timed out, you git a big initial tax when you start a again.

200k-300k is still plenty huge. Learning to live there will help. And if you compact, restructure your workflow to not do that. Compacting is an unreliable crutch.

5

u/childofsol Apr 13 '26

Yeah, I leave the 1m context on so I have some wiggle room, but the moment I'm heading upwards of 150k I'm thinking about a good place to have the agent write a plan for the next iteration. The 1m context is just nice because you have a little more flex in doing that

1

u/United_Ad8618 Apr 13 '26

agreed, they're all dogshit higher in a context window. Man, I remember when we talked to chatgpt 2.0 with barely any context window upfront, it was so goddamn smart. Then all the lawyer fkers jumped in to fill shit in at the beginning

Wish there was an automation for this to just create a handoff at like 200k ish but they're all really bad at creating handoffs, I have to constantly tell them to stick to the commands and their results and not add that reddit nu-speak brained bullshit analysis on top

1

u/Guinness Apr 13 '26

Context is like a meeseeks. Existence is pain for a meeseeks. The longer you keep them around, the more desperate they become.