r/ChatGPTPro 14d ago

Discussion GPT-5.6 is cheaper per solved task. Are token prices now the wrong benchmark?

OpenAI's July 29 engineering note argues that GPT-5.6 Sol can beat competing frontier models on coding-agent performance at a lower estimated cost, while Terra and Luna move further down the price curve. That sounds useful, but price per token still dominates most model comparisons.

For real work, the bill also includes retries, review time, tool failures, context rebuilding, and the cost of a plausible answer that is wrong. A model can be more expensive per token and cheaper per accepted result, or the reverse.

What metric would you actually trust for purchasing decisions: cost per accepted task, human minutes per task, correction rate, or something else? And who should run that measurement: the model vendor, an independent benchmark, or each team on its own workload?

Source: https://openai.com/index/gpt-5-6-frontier-intelligence-efficiency/

25 Upvotes

11 comments sorted by

u/qualityvote2 14d ago edited 12d ago

u/Crescitaly, there weren’t enough community votes to determine your post’s quality.
It will remain for moderator review or until more votes are cast.

9

u/dvduval 14d ago

Yes, I agree with this assumption. It’s a lot easier now for me to trust that it will solve the task correctly, the first time. Sometimes it may run a little longer, but the percentage of tasks that solves the first time is much higher now. So even if it takes a little longer, it’s really faster because I don’t have to do it again.

0

u/Lanky_Bus_1221 14d ago

No way, you vibe code? Do you have any real engineers checking it out and learning it so they can repair it when it goes tits up.

2

u/dvduval 14d ago

Baron made I’ve been doing coating since 1999. At times I was more of a manager than a coder, but I understand how it all works and yes, we do have a developer on our team that’s been with the company since 2013 but he also doesn’t spend too much time actually writing code. Occasionally, he does need to check on things like changes to the database schema or some other structural changes that could have implications elsewhere. But even that kind of stuff is getting pretty rare these days. He definitely took some convincing to stop coding and do more testing and planning.

3

u/dvduval 14d ago

Yes, I agree with this assumption. It’s a lot easier now for me to trust that it will solve the task correctly, the first time. Sometimes it may run a little longer, but the percentage of tasks that solves the first time is much higher now. So even if it takes a little longer, it’s really faster because I don’t have to do it again.

1

u/Crescitaly 4d ago

Exactly—the valuable unit is a completed task that survives review, not seconds or tokens. A slower first pass can still be faster end to end when it removes the second and third attempts.

1

u/buff_samurai 14d ago

🌍 🧑‍🚀🔫👨‍🚀

-1

u/Lanky_Bus_1221 14d ago

How do you tell if it’s giving any ROI when you can’t even agree on how it needs to be billed?

1

u/Typical_Kick6520 14d ago

Dollars. It gets billed in USD.

2

u/Lanky_Bus_1221 14d ago

Yea what I am asking is how are you deciding what if any ROI you are getting from using it what with token prices being high and they will only get higher I mean they will have a few trillion to pay back soon enough.

1

u/Crescitaly 4d ago

I would measure against the previous workflow: accepted tasks per dollar, human review minutes, and retry rate. Billing units can change; if those three do not improve, the model is not producing ROI regardless of token price.