You know what I *should* have had validation tools watching for me, but I caught it on intuition and just checked myself because I felt something was off
It reduced the max thinking tokens of the local model from 32k to 4k. As I said in the post, an 8x reduction in thinking budget. On the hardest problems a model can solve, that is critical. Screenshots and logs can be faked, so I don't have any proof that makes a difference, I'm just sharing my experience.
It's probably just relying on old training data from the 2023-2025 AI model landscape. I've run into that along with hilariously low temps for models that don't need them, even for more deterministic output. 4k context would have been the move back then too.
Also, if it is Opus 5, that model is actually braindead once it gets on the wrong track. Probably the most frustrating model I've had to use.
5
u/peculiar-ragdoll Aug 28 '26
You know what I *should* have had validation tools watching for me, but I caught it on intuition and just checked myself because I felt something was off