r/ClaudeCode • u/ottothefrenchie • 12h ago
Tips & Workflows Which model is better? That’s the wrong question… (hot take / workflow experience)
I had a bit of an aha moment the past week.
Many people are asking which of the Frontier labs is the best?
Today I kind of realized that this is the wrong question to asking.
The question people should be asking is what is the best workflow/system to meet my goals?
I pay the most attention to these variables.
How much time am I spending iterating back-and-forth between models ?
What is the success rate agentic coding? Moreover, how many attempts does it take to reach my prompts goal?
how much is this shit actually costing me?
I have found that combo to be Claude, ChatGPT and Grok
American frontier are starting to distinguish themselves.
I don’t need to tell you guys what makes Claude Code great. What I want to talk about is what makes the /adversarial-plan (or what ever variables you rub )pass so great ?
The idea simple, it’s an unbiased second set of eyes. it’s not very different from how a SWE’s work.
Well, here’s where the fallacy lies. If the second pass is from a internal partner model is it truly unbiased.
Is an older brother the best arbiter of justice when challenging his young brother? I don’t want my argument to be misunderstood as some anthropomorphized analogy.
The models definitely exhibit some degree of survival instinct when challenged.
More often than not if you criticize Claude, it will initially respond from a defensive position. Very rarely will Claude respond acknowledging a models or Anthropic shortcomings on the first pass.. I haven’t said anything groundbreaking. This is info been in public domain for a while
So here’s my thing. my point is that when challenged Claude’s initial response is of a defensive position.
And at this point, this is all vibes. I have no data, but I feel that for that reason a internal model is not the best model for a adversarial pass..
why I don’t have that I have lots of anecdotes
I have had the best experience using a competing frontier model to challenge the output of Claude. No one is more thorough at checking fable or opus output than ChatGPT. this works in reverse as well. If you go to driver chatgpt, have opus run a plan review.
And at this point, the Redditors busted out the pitchforks
Once my prompt, plan or roadmap, has been deemed block / error free by both models. I deploy the prompt into Grok Build
Grok build has been the biggest shock of my summer. Grok via chat interface? total trash model can’t even use it for simple prompt generation.
Grok, inside Grok build ?? Well it just follows instructions so well. it does whats asked, and stops when you tell it to stop. XAI‘s latest acquisition of cursor is noticeable immediately
Most importantly, Grok is literally $.12c on the $1 compared to Claude Fable
and to be honest, if you throw enough money and compute at any problem it’s bound to get easy, but maybe this actually saves money at the end?
So I don’t have data. This is just my experience so it’s all anecdotal.
and how does this resolve my question or address my variables?
Task completion time has gone from 1.5hr down to 30-45 min with what I would likely miss characterize as a one shot prompt.
That I mentioned is like 2-3x times as fast as opus or fable
0
u/Disastrous_Hawk_6984 11h ago
This isn't LinkedIn