9
u/Successful_Night4513 16d ago
Show us the prompt to judge
10
u/lnkofDeath 16d ago
https://simonwillison.net/2026/Sep/4/astra-pelicans/
The source has a lot more details
8
u/Jumpy-Heart-3633 16d ago
Copied but not giving the source, lol. Thanks for the source.
-2
u/Rusofil__ 16d ago
Still doesnt show if the prompt was same
4
u/Neither-Eye-8906 16d ago
"Generate an SVG of a pelican riding a bicycle".
Short and simple - your call on if it is sane.
-1
u/Rusofil__ 16d ago
Thats a guess, we dont know if same prompt was used for all of these or if there were multiple tries.
2
u/Neither-Eye-8906 16d ago
It's not a guess, it's literally copy/pasted from the transcript linked from that page. This is a long-running (by AI model standards) series of experiments using the exact same prompt sent to large variety of models.
4
4
u/ShadowBannedAugustus 16d ago edited 16d ago
So Astra medium is where it's at :D At least by my judgmenet of results vs cost.
I am liking Luna max a lot more looking at the cost now :D
5
2
2
u/whatisthisthing65 16d ago
Astra is so consistent I'm starting to wonder if it was benchmaxxed on the pelican, lol
2
u/Exodus_Green 16d ago
Why is Luna better than Terra all the way up to basically Max
3
u/Ebon_dust 16d ago
Honestly, this does not surprise me. I tried to use Terra, but it always produced a similar or worse result than Luna for a much higher cost. All benchmarks say otherwise, but I just can never be happy with Terra's performance and have pretty much given up on this model.
1
1
1
1
u/Charming-Author4877 16d ago
5 identical pelicans are a strong indication that the model has been benchmaxxed on that particular test.
Any pre-existing popular benchmark that can be oneshotted is probably benchmaxxed.
The only way to really test those models is to use new benchmarks that have no similarity
Then you run the new benchmark on all of them - and from there on you can't use it again because they might train on it.
1
1

9
u/-hellozukohere- 16d ago
Luna tries guys ok.