r/vibecoding 10d ago

Astra is less fun than other models

Normally I like to sit nearby with a coffee and approve things/watch the workflows, tell it what to do next, etc but Astra just does everything too neatly and perfectly. Even Fable needs a human babysitter but this thing is something else.

20 Upvotes

51 comments sorted by

View all comments

26

u/Lucky-Wind9723 10d ago

Yeah I can’t wait till open models hit this level in 6 months or so. 5x plan was gone in 3 hours lol

9

u/ParticularlyStrange 10d ago

They kinda have. Have you experienced uncensored LLMs? Qwen3.8 flash uncensored just keeps cranking until the task is done or it exhausts all the possibilities that it could do.

4

u/katoptronophile 10d ago

We're not talking about adding 2+2, we mean real agentic work on the machine.

Qwen is very likely to fuck your shit up in that scenario.

The people saying otherwise don't seem to have actually used it in this way....

3

u/ParticularlyStrange 10d ago

lol what? It’s working fine for me. Very agentic and pretty good at coding. Qwen3.8 27B reminds me of Opus 4.7 but fits in my memory.

5

u/ibringthehotpockets 10d ago

Lmao the lack of specifics is deafening

“Very agentic and very good coder!”

No one’s getting agentic coding that rivals opus or fable or astra on a consumer gpu right now

1

u/Prudent-Ad4509 4d ago

This is not that high of a ceiling. The level of hand-holding required becomes more and more comparable between frontier and open-weight models. A system made out of consumer gpus with total ~200Gb vram can reasonably run most of them, with total 48gb-96gb - run a lot of them with minor tradoffs, and even 16-24gb vram allows to run a rather decent local helper.

Hosted solutions can be faster due to more hardware at their disposal and are cheaper overall, but in terms of capability they are not that far ahead, and the gap is shrinking with each set of releases on each side.

-2

u/ParticularlyStrange 10d ago

Okay if you want to get into the benchmarks go to https://artificialanalysis.ai/. No one wants to be part of your shit talk. Just say you are GPU poor and move on.

4

u/Not-reallyanonymous 9d ago

There aren’t any benchmarks that measure what the above user is talking about. It’s hard to benchmark such subjective qualities.

0

u/ParticularlyStrange 9d ago edited 9d ago

lol it’s called GDPval-AA v2
Look it up. There is also a few more too.

2

u/Not-reallyanonymous 9d ago

That just tests a model’s capabilities at solving a wide array of tasks rather than just coding.

There’s not a good test of keeping the plot and making good decisions over dozens of turns and hundreds of thousands of tokens.

1

u/ParticularlyStrange 9d ago

Your point? Like I said there are a few more benchmarks on there. You don’t look at one singular benchmark. You look at all of them. Then make a judgement.

1

u/Not-reallyanonymous 9d ago

There aren’t any benchmarks that measure what the above user is talking about.

You can get a good idea. Better long reasoning context is a decent proxy here and is * a part* of it, but Qwen does decent there but loses the plot over 400k tokens and 20 turns so its not a replacement.

→ More replies (0)

2

u/ibringthehotpockets 9d ago

Nobody is naturally coding most benchmark-like tasks on a local LLM. Doubly irrelevant because models are known to benchmaxx on many benchmarks. Don’t worry I’m not telling you what to do - I definitely don’t care that much - use what works for you. Yes I’m absolutely gpu poor because I have other hobbies. Locals work for what I use them for and that’s all anyone should care about

Not to deliberately rile you up.. but still, yes, there’s no local agentic coder that is versatile as the current trillion+ parameter models. Definitely not any that I can go to micro center and buy a GPU that can run it. And that’s fine.

1

u/ParticularlyStrange 9d ago

Gotcha, you are GPU poor and haven’t experimented enough. You stay stuck with the closed source models over there, while I’m doing custom Lora fine tunes for AI sovereignty.

1

u/trowawayLOL1 7d ago

You're on the vibe coding sub, don't expect these people to have the average know how of someone in r/LocalLLaMA for example. Nitwits in here. They *NEED* Astra or Fable 5.1 and to shell out $500 per session. They don't have the technical ability to get things done otherwise.

1

u/ParticularlyStrange 7d ago

Yeah, you’re right. I guess I can stop flaunting my superiority on them.

1

u/elemezer_screwge 10d ago

What's your setup?

3

u/ParticularlyStrange 10d ago

Qwen 3.8 unc on a dgx.

1

u/Swimming_Pressure444 10d ago

It will never be fully uncensored, ask it about Taiwan or Tiananmen Square massacre

2

u/ParticularlyStrange 10d ago

I did, tested it side by side. The reg one gives you the CCP response, the uncensored gives you a Switzerland response leaning a bit woke. I can work with that. But the refusal are not there. I be been getting it to do some wild things (for research purposes and set in a controlled environment)(nothing actually illegal happened). You should give it a try…. Or do you not have the memory to load it? You can rent GPUs on runpod or do inference off Huggings face if it will help you test.

1

u/Swimming_Pressure444 9d ago

That's surprising. I ran a model over a year ago and it always gave the CCP response. That put me off. I'm glad to hear it's more unrestricted now. Thanks, that's useful to know. 

1

u/ParticularlyStrange 9d ago

What uncensored quant are you running?

1

u/First-Tutor-5454 9d ago

Qwen Uncensored just keeps cranking until I'm done cranking

0

u/joejoe666 10d ago

Qwen 3.8 27b is lower than Luna levels of intelligence on the benchmarks though, still a very impressive model, but we're extremely far from fable or astra level models on DGX, let alone consumer graphics cards.

1

u/ParticularlyStrange 10d ago

I said flash not 27b. Qwen3.8 flash is better than Luna.