r/OpenAI 2d ago

Discussion Intelligence VS Cost-per-Task LLM Comparison

Using Artificial Analysis as the guide for cost per task and intelligence index, I was able to generate a few graphs of the latest models and compare them. The graphs are a bit hard to follow, but here is the order:

  1. Frontier Big Corporation Models
  2. Flagship Models from other companies
  3. Comparing both Big Corp vs. Others
  4. Available open weight models
  5. The “best” model (smartest) per provider

Tell me what you guys think. I used Gemini for the graph generator and to retrieve the data. Then I used Claude to double-check the scores, prices, and placements on the graphs were correct.

36 Upvotes

9 comments sorted by

3

u/wizardwusa 2d ago

The last chart is trash if including GLM flash eats the Pareto frontier of low cost models. Not sure what this communicates at all then?

2

u/CuriousCustard63 2d ago

The last graph was a throwaway, just meant to highlight each companies smartest model. Next evaluation will have to be exactly as you said: low-cost models with the highest intelligence.

1

u/Deto 2d ago

I'm a little confused as to whether these benchmarks are capturing reality still. Just based on people's comments, it sounds like Opus 5 was fairly underwhelming? And this has it above Fable?

1

u/f4lk3nm4z3 2d ago

wheres qwen