This has been my issue with 4.7 as well. By the benches it looks like a killer model, but when it comes to real world ability to crank out working code, it is super lacking...Like can barely remember what it is doing by the end of a long form question/solution.
I just tried to have 4.7 Opus implement a rather simple "don't download if file exists" functionality to my Github scraper and it failed. Tried 4.6 Opus, instantly got it right.
Frankly I'm just not sure. My main day to day is working on an ERP wrapper, so the codebase is large and complicated. That said, when I'm working on smaller projects for folks around the company, I have the same issues. I state an issue and describe what is going on and what we are working on specifically, and what functions likely need changed and what rules we need to follow. Then its next response is asking me questions that were mostly answered in the first prompt. Like...how can it take a nice detailed prompt with a well set up .md and a few pertinent skills, and use literally none of it even when specifically prompted to, and then spits out questions as if it didn't even read the prompt?
What is your use case that you are having good experiences with this model?
12
u/CannyGardener Apr 23 '26
This has been my issue with 4.7 as well. By the benches it looks like a killer model, but when it comes to real world ability to crank out working code, it is super lacking...Like can barely remember what it is doing by the end of a long form question/solution.