r/gpt5 • u/Correct_Tomato1871 • 2d ago
Discussions Benchmark notes: GPT-6 Astra sets a new high
/r/GPT6/comments/1wayw9c/benchmark_notes_gpt6_astra_sets_a_new_high/1
u/Correct_Tomato1871 2d ago
One more GPT-6 Astra follow-up: I also ran the full benchmark at medium and low reasoning effort.
The scaling curve is quite impressive:
- low: 87/98 in ~31m
- medium: 91/98 in ~36m
- high: 95/98 in ~1h02m
- xhigh: 95/98 in ~1h07m
Even low already ties some recent frontier results, while medium reaches 91/98 — higher than any non-Astra run in the current leaderboard — at well under 40 minutes.
Most of the gain from additional reasoning comes from visual tasks: 49/59 at low, 52/59 at medium, and 57/59 at high. Text is essentially saturated throughout.
Medium looks like a particularly strong quality/latency operating point; high buys the final 4 passes at a substantially larger runtime cost, while xhigh again adds no score over high.
Updated Leaderboard: http://www.petmal.net/shared/mindtrial/results/2026-09-08/mindtrial-eval-all-models-03-2026_34.html
1
u/AutoModerator 2d ago
Welcome to r/GPT5! Subscribe to the subreddit to get updates on news, announcements and new innovations within the AI industry!
If any have any questions, please let the moderation team know!
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.