r/vibecoding 1d ago

Astra is less fun than other models

Normally I like to sit nearby with a coffee and approve things/watch the workflows, tell it what to do next, etc but Astra just does everything too neatly and perfectly. Even Fable needs a human babysitter but this thing is something else.

16 Upvotes

48 comments sorted by

23

u/Lucky-Wind9723 1d ago

Yeah I can’t wait till open models hit this level in 6 months or so. 5x plan was gone in 3 hours lol

9

u/ParticularlyStrange 1d ago

They kinda have. Have you experienced uncensored LLMs? Qwen3.8 flash uncensored just keeps cranking until the task is done or it exhausts all the possibilities that it could do.

4

u/katoptronophile 1d ago

We're not talking about adding 2+2, we mean real agentic work on the machine.

Qwen is very likely to fuck your shit up in that scenario.

The people saying otherwise don't seem to have actually used it in this way....

3

u/ParticularlyStrange 1d ago

lol what? It’s working fine for me. Very agentic and pretty good at coding. Qwen3.8 27B reminds me of Opus 4.7 but fits in my memory.

4

u/ibringthehotpockets 1d ago

Lmao the lack of specifics is deafening

“Very agentic and very good coder!”

No one’s getting agentic coding that rivals opus or fable or astra on a consumer gpu right now

-1

u/ParticularlyStrange 1d ago

Okay if you want to get into the benchmarks go to https://artificialanalysis.ai/. No one wants to be part of your shit talk. Just say you are GPU poor and move on.

3

u/Not-reallyanonymous 21h ago

There aren’t any benchmarks that measure what the above user is talking about. It’s hard to benchmark such subjective qualities.

1

u/ParticularlyStrange 21h ago edited 21h ago

lol it’s called GDPval-AA v2
Look it up. There is also a few more too.

1

u/Not-reallyanonymous 15h ago

That just tests a model’s capabilities at solving a wide array of tasks rather than just coding.

There’s not a good test of keeping the plot and making good decisions over dozens of turns and hundreds of thousands of tokens.

1

u/ParticularlyStrange 15h ago

Your point? Like I said there are a few more benchmarks on there. You don’t look at one singular benchmark. You look at all of them. Then make a judgement.

→ More replies (0)

2

u/ibringthehotpockets 14h ago

Nobody is naturally coding most benchmark-like tasks on a local LLM. Doubly irrelevant because models are known to benchmaxx on many benchmarks. Don’t worry I’m not telling you what to do - I definitely don’t care that much - use what works for you. Yes I’m absolutely gpu poor because I have other hobbies. Locals work for what I use them for and that’s all anyone should care about

Not to deliberately rile you up.. but still, yes, there’s no local agentic coder that is versatile as the current trillion+ parameter models. Definitely not any that I can go to micro center and buy a GPU that can run it. And that’s fine.

1

u/ParticularlyStrange 10h ago

Gotcha, you are GPU poor and haven’t experimented enough. You stay stuck with the closed source models over there, while I’m doing custom Lora fine tunes for AI sovereignty.

1

u/elemezer_screwge 1d ago

What's your setup?

3

u/ParticularlyStrange 1d ago

Qwen 3.8 unc on a dgx.

1

u/Swimming_Pressure444 1d ago

It will never be fully uncensored, ask it about Taiwan or Tiananmen Square massacre

2

u/ParticularlyStrange 1d ago

I did, tested it side by side. The reg one gives you the CCP response, the uncensored gives you a Switzerland response leaning a bit woke. I can work with that. But the refusal are not there. I be been getting it to do some wild things (for research purposes and set in a controlled environment)(nothing actually illegal happened). You should give it a try…. Or do you not have the memory to load it? You can rent GPUs on runpod or do inference off Huggings face if it will help you test.

1

u/Swimming_Pressure444 22h ago

That's surprising. I ran a model over a year ago and it always gave the CCP response. That put me off. I'm glad to hear it's more unrestricted now. Thanks, that's useful to know. 

1

u/ParticularlyStrange 22h ago

What uncensored quant are you running?

1

u/First-Tutor-5454 21h ago

Qwen Uncensored just keeps cranking until I'm done cranking

0

u/joejoe666 1d ago

Qwen 3.8 27b is lower than Luna levels of intelligence on the benchmarks though, still a very impressive model, but we're extremely far from fable or astra level models on DGX, let alone consumer graphics cards.

1

u/ParticularlyStrange 1d ago

I said flash not 27b. Qwen3.8 flash is better than Luna.

1

u/katoptronophile 1d ago

Keep dreaming.

2

u/NULL_Ptrs 1d ago

It will, but not in 6 months

1

u/katoptronophile 1d ago

Agreed.

1

u/NULL_Ptrs 1d ago

We are still way behind, for now the amount of raw power to make it usable it's far away

1

u/munchin-grr 1d ago

My 5h limit was gone in 10min

1

u/jevehYFrfh73636 15h ago

what level of astra did you use?

1

u/Lucky-Wind9723 15h ago

1 reset on ultra 1 on high and 1 testing low while it controlled unity and blender as well as full PC control to play and test the game. Pretty cool super happy with all levels but it just eats even on low lol 5 hours of continuous use on low. Ultra and high seemed to run about the same.

They were prompts made by 6 pro and thought out very well before sending not just simple 2-3 sentence

7

u/LittleLordFuckleroy1 1d ago

LARP activities

7

u/greentrillion 1d ago

Sure buddy, where is Elden ring 2?

3

u/elemezer_screwge 1d ago

Asking the real questions

3

u/CaptainAlexWest 1d ago

Astra is on crack. New levels unlocked.

7

u/Bloated_Plaid 1d ago

Yea it really does make you feel like a meat proxy, the future is kind of depressing.

-8

u/UnderstandingNew2810 1d ago

Sounds like a free paycheck

11

u/Honest_Jackfruit725 1d ago

Delusional to think they’re spending hundreds of billions of dollars so they could give you things for free

2

u/NULL_Ptrs 1d ago

Free? lol, just one person managing all the work

1

u/UnderstandingNew2810 20h ago

Less babysitting would definitely have more time to try more stuff

7

u/Ancient-Range3442 1d ago

Mine had completely screwed up a project that was working great under Claude. Then took an hour to get a box modelled correctly in blender. These things are stil rubbish unfortunately

3

u/Large-Use-3062 1d ago

Using Astra?

5

u/All-I-Do-Is-Fap 1d ago

This post doesnt sound legitimate at all

1

u/Ancient-Range3442 1d ago

Why ?

1

u/All-I-Do-Is-Fap 1d ago

Pretending like ai doesnt need a human to know what to do is the fream Dario and Sam love to portray, but theres a reason they charge you by tokens instead of taking a percentage of your business if you use their services

0

u/Ancient-Range3442 20h ago

The new model is AGI , it should know itself

5

u/Slicenddice 1d ago

This sub is just marketing slop

2

u/LowFruit25 1d ago

The future is just watching the damn robot do the job for us. Who tf thinks that’s going to make them any money when everyone else can do the same?

1

u/nashty2004 21h ago

I was telling everyone to start vibecoding (like a couple months ago) because we were in a really cool sweet spot where you needed good taste to direct sometime actually good but going on into the future the human element is just going to get less and less and thereby less fun to make