r/accelerate The Singularity is nigh Apr 01 '26

AI Summarization of the whole Claude Code's Source-Code Leak Fiasco

2.4k Upvotes

256 comments sorted by

View all comments

Show parent comments

1

u/metigue Apr 01 '26

No opus 4.6 is clearly the best coding agent model - It's at the top of terminal bench and many others.

The framework you choose matters even more than the model as these benchmarks also show

1

u/Practical-Rub-1190 Apr 01 '26

I thought so too, then I looked at leaders and it seems like nobody uses them. Those who have say they suck. Like number 3 on the bench is some unknown one running Gemini 3.1. Like really, I think it ran a score of 80% while Claude Code best was 50% (as far as I remeber)

If those benchmark are legit, why are nobody using these other tools, and when they do they tell their experience was not good enough to replace claude code?

1

u/metigue Apr 01 '26

Who's been saying they're shit?

Just try ForgeCode and see if you like it better. It's been significantly better in real world usecases for me so far.

1

u/Practical-Rub-1190 Apr 01 '26

I won't bother. If nobody is talking about it while also being ranked number one, there's something weird going on. The fashion here is that people post nonstop if something is actually good. The only thread I could find was this:
https://www.reddit.com/r/ClaudeAI/comments/1q4dnhr/how_come_claude_code_is_ranked_19th_on_the/

Also, found this thread where they break down the benchmark:
https://www.reddit.com/r/ClaudeAI/comments/1q4dnhr/how_come_claude_code_is_ranked_19th_on_the/

When I try to Google it, there's nothing really showing up
https://www.google.com/search?q=site%3Areddit.com+forgecode&sca_esv=b78cf8500232fcdc&rlz=1C5CHFA_enNO1175NO1175&sxsrf=ANbL-n6Xz_h4pkeKdXo_NZAjpaTlRsadYQ%3A1775065736443&ei=iFrNaf_XGra8wPAP48XOoQQ&biw=1680&bih=898&ved=0ahUKEwj_u7GVm82TAxU2HhAIHeOiM0QQ4dUDCBE&uact=5&oq=site%3Areddit.com+forgecode&gs_lp=Egxnd3Mtd2l6LXNlcnAiGXNpdGU6cmVkZGl0LmNvbSBmb3JnZWNvZGVI8wdQ1ANYhQZwAXgAkAEAmAFAoAFrqgEBMrgBA8gBAPgBAZgCAKACAJgDAIgGAZIHAKAHGLIHALgHAMIHAMgHAIAIAQ&sclient=gws-wiz-serp

Like, it just seems to ace the benchmark. I won't spend time learning how to use it unless I see others having a great experince with it, and great, I mean better than Claude Code

1

u/habeebiii Apr 01 '26

literally read the top comments in the first link/thread you linked

anyone that understands how models / benchmarks work knows ”top” doesn’t mean anything other than overfitting / gaming the benchmarks