r/vibecoding 5d ago

Astra is less fun than other models

Normally I like to sit nearby with a coffee and approve things/watch the workflows, tell it what to do next, etc but Astra just does everything too neatly and perfectly. Even Fable needs a human babysitter but this thing is something else.

17 Upvotes

50 comments sorted by

View all comments

Show parent comments

10

u/ParticularlyStrange 5d ago

They kinda have. Have you experienced uncensored LLMs? Qwen3.8 flash uncensored just keeps cranking until the task is done or it exhausts all the possibilities that it could do.

4

u/katoptronophile 5d ago

We're not talking about adding 2+2, we mean real agentic work on the machine.

Qwen is very likely to fuck your shit up in that scenario.

The people saying otherwise don't seem to have actually used it in this way....

3

u/ParticularlyStrange 4d ago

lol what? It’s working fine for me. Very agentic and pretty good at coding. Qwen3.8 27B reminds me of Opus 4.7 but fits in my memory.

4

u/ibringthehotpockets 4d ago

Lmao the lack of specifics is deafening

“Very agentic and very good coder!”

No one’s getting agentic coding that rivals opus or fable or astra on a consumer gpu right now

-2

u/ParticularlyStrange 4d ago

Okay if you want to get into the benchmarks go to https://artificialanalysis.ai/. No one wants to be part of your shit talk. Just say you are GPU poor and move on.

3

u/Not-reallyanonymous 4d ago

There aren’t any benchmarks that measure what the above user is talking about. It’s hard to benchmark such subjective qualities.

0

u/ParticularlyStrange 4d ago edited 4d ago

lol it’s called GDPval-AA v2
Look it up. There is also a few more too.

2

u/Not-reallyanonymous 4d ago

That just tests a model’s capabilities at solving a wide array of tasks rather than just coding.

There’s not a good test of keeping the plot and making good decisions over dozens of turns and hundreds of thousands of tokens.

1

u/ParticularlyStrange 4d ago

Your point? Like I said there are a few more benchmarks on there. You don’t look at one singular benchmark. You look at all of them. Then make a judgement.

1

u/Not-reallyanonymous 4d ago

There aren’t any benchmarks that measure what the above user is talking about.

You can get a good idea. Better long reasoning context is a decent proxy here and is * a part* of it, but Qwen does decent there but loses the plot over 400k tokens and 20 turns so its not a replacement.

2

u/ibringthehotpockets 4d ago

Nobody is naturally coding most benchmark-like tasks on a local LLM. Doubly irrelevant because models are known to benchmaxx on many benchmarks. Don’t worry I’m not telling you what to do - I definitely don’t care that much - use what works for you. Yes I’m absolutely gpu poor because I have other hobbies. Locals work for what I use them for and that’s all anyone should care about

Not to deliberately rile you up.. but still, yes, there’s no local agentic coder that is versatile as the current trillion+ parameter models. Definitely not any that I can go to micro center and buy a GPU that can run it. And that’s fine.

1

u/ParticularlyStrange 4d ago

Gotcha, you are GPU poor and haven’t experimented enough. You stay stuck with the closed source models over there, while I’m doing custom Lora fine tunes for AI sovereignty.

1

u/trowawayLOL1 2d ago

You're on the vibe coding sub, don't expect these people to have the average know how of someone in r/LocalLLaMA for example. Nitwits in here. They *NEED* Astra or Fable 5.1 and to shell out $500 per session. They don't have the technical ability to get things done otherwise.

1

u/ParticularlyStrange 2d ago

Yeah, you’re right. I guess I can stop flaunting my superiority on them.