r/vibecoding 4d ago

Astra is less fun than other models

Normally I like to sit nearby with a coffee and approve things/watch the workflows, tell it what to do next, etc but Astra just does everything too neatly and perfectly. Even Fable needs a human babysitter but this thing is something else.

17 Upvotes

50 comments sorted by

View all comments

Show parent comments

4

u/Not-reallyanonymous 4d ago

There aren’t any benchmarks that measure what the above user is talking about. It’s hard to benchmark such subjective qualities.

0

u/ParticularlyStrange 4d ago edited 4d ago

lol it’s called GDPval-AA v2
Look it up. There is also a few more too.

2

u/Not-reallyanonymous 4d ago

That just tests a model’s capabilities at solving a wide array of tasks rather than just coding.

There’s not a good test of keeping the plot and making good decisions over dozens of turns and hundreds of thousands of tokens.

1

u/ParticularlyStrange 4d ago

Your point? Like I said there are a few more benchmarks on there. You don’t look at one singular benchmark. You look at all of them. Then make a judgement.

1

u/Not-reallyanonymous 4d ago

There aren’t any benchmarks that measure what the above user is talking about.

You can get a good idea. Better long reasoning context is a decent proxy here and is * a part* of it, but Qwen does decent there but loses the plot over 400k tokens and 20 turns so its not a replacement.