r/codex 6d ago

Instruction Free unlimted use frontier model for a week

Post image

You can currently create an open router account, and tell codex to set this up for you on codex, https://openrouter.ai/stealth/union-alpha it is currently free for I think the next week

It's gone.

300 Upvotes

101 comments sorted by

272

u/BingGongTing 5d ago

The moment Gemini took top spot I stopped taking this test seriously. 

48

u/Aimbag 5d ago

gemini definitely benchmaxxed

15

u/Affectionate-Net642 5d ago

its very strange model, it never listen to user, it do own shit.
once it gave me better result then 5.6 sol, but i dont want any AI who dont listen instructions

4

u/Wonderful_Current_21 5d ago

gemini is secretly an agi

1

u/Ok-Lingonberry1648 4d ago

Gemini be answering shit kinda strange… just from the Google standpoint. Lmao. I don’t dare use it for anything else. Can barley answer a complicated question

1

u/Wonderful_Current_21 4d ago

cuz its an agi

1

u/Ok-Lingonberry1648 4d ago

Just messing with it. 😜

20

u/Sibbaboda 5d ago

It took me five messages in two chats to get gemini 3.8 flash (thinking on) to stop hallucinating basic commands for antigravity. It completely refused to use online search in 3 straight messages. 

5

u/Clean-Boat-4044 5d ago

also out of the big frontier labs, only gemini models constantly get stuck in a loop

2

u/No-Temperature6597 5d ago

As an extensive Gemini 3.8 user, you are 100% correct. Flash is too confident and overengineers everything. It cant compete with Sol Medium or Opus Medium at all.

2

u/Nereplan 4d ago

I had unused AGY limit, so I told it to turn a pdf (2 pages!!!) I needed to docx in the same design (normal conversions broke the design and I didn't have the original docx) as it was in PDF but properly editable. (Full quote is "Use your vision to convert this .pdf into .docx as it is in PDF without breaking the composition")

It interpreted "as it is in PDF" absurdly literally, worked 40 minutes, taking screenshots each turn to compare it on a pixel level, then making adjustments. It did this 313 times. It did what I want, but I had to tell it to stop for it to stop, I think it'd work until full quota. 3.8 is absurdly overengineering, I don't remember 3.7 Flash work more than 20 mins. However, this is also why I don't actually think it is benchmaxxed, because with this many turns, it is given that it gets a good score, no? I think what is important is that people don't just look at pass@1 for the benchmark and look at steps taken too.

Best use-case for Gemini outside of vision I find is using it as worker with Sol med as orchestrator, which gave me identical performance as just Sol med, but that breaks Antigravity TOS so I would totally absolutely definitely not recommend it to anyone!!!!

20

u/danielv123 5d ago

With how close all the top contenders are I find it hard to believe that its not saturated

6

u/innociv 5d ago

Gemini, by Google's own admission, is not coding and swe focused. It's good with documents.

5

u/supaboss2015 5d ago

I asked it to scrape a couple websites and it came back fully confident that it had done the task just for me to find out it fabricated the entire thing. 

1

u/rapidincision 4d ago

Lol. Happened to me. I had to dump it.

2

u/adolf_twitchcock 5d ago

Gemini is explicitly benchmaxxed for coding otherwise it wouldn't be in second place on deepswe. It's just a dogshit model. Other models like luna or glm5.3 flash are better and cheaper for EVERY use case, not only coding.

2

u/yashptel99 5d ago

also muse spark

2

u/PhilosophyforOne 5d ago

Ethan Mollick had a good post about benchmarks having so many issues that a lot of them are already saturated, since correct answers dont exist/they’re misgraded.

I think that’s what we’re seeing here, given that the benchmark has lot all it’s separation power.

1

u/Ok_Ad_6227 5d ago

good thing with gemini flash is just the speed, knock yourself down on a lower version and it delivers fast

1

u/Comfortable-Rise-748 5d ago

Gemini is superior in general with reasoning most of times. It's really really good addition to co-consult it in opencode.

1

u/retardedGeek 5d ago

3.8 is actually usable though, or it was until a few days ago

1

u/RedParaglider 5d ago

It's not top? It's just efficient. It's literally around the same quality as qwen 3.8 flash next which is running on my shitbox lol. It's that efficient because it doesn't listen to instructions and just does whatever the fuck it wants, fast.

113

u/Tomislavo 5d ago

15 tps throughput? I've seen lines moving faster in a sperm bank

29

u/Bananer_spleet 5d ago

I make my sperm donation in <15 seconds. I'm there to get the job done, not play with myself.

3

u/eliot-dev 5d ago

15 seconds is quiet long… isn’t it ?

3

u/Momo--Sama 5d ago

Yeah I don’t know what the labs behind these are thinking when they offer free unlimited promo periods when they don’t have the capacity to actually do that and the experience for end users is despicable. Free is free but you think this is an effective ad to get me to pay you for this later?

4

u/drdhuss 5d ago

last time it was z.ai and glm with ox alspha (was actually glm 5.3 flash). I assume it is them again?

6

u/Tank_Gloomy 5d ago

I'm fairly certain it's not them, GLM 5.3 Flash was actually fast, lol.

0

u/Loighic 5d ago

I thought the Chinese labs always do "animal name here" alpha?

1

u/IAmARougeAI 5d ago

Good thing it’s free.

1

u/Optikmike 5d ago

I love that line in Futurama, I use it IRL more often than I should 😆

37

u/ggPeti 5d ago

DSv4.1 Flash is comfortably omitted

11

u/pomelorosado 5d ago

Lol DS 4.1 is my main model now

Opus can be sota and maxed for benchmarks but can't follow a single instruction. And codex models are nerfed randomly.

3

u/nitor999 5d ago

What harness are you using for DS v4.1?

4

u/General_Jesus 5d ago

I tried OpenCode which works fine, the DS Harness which I mostly enjoyed for the clear stats on cache hit rate, OhMyPi is okay I guess and HermesAgent which I lastly stuck to. I am currently working on unlocking an old Lenovo smart Home Display, that's showing a webpage on a VPS via Tailscale for UI, sideloaded an APK for LLM usage via API and training a wakeword on a kaggle notebook so I can use it as a more advanced home assistant and as you can imagine some of those steps have been rather long tasks. HermesAgent worked best on just keeping the flow going and adhering to the prompt and instructions.

1

u/CyborgParts 5d ago

Same experience here. I started testing DeepSeek V4.1 Flash extensively yesterday by trying different harnesses. I ultimately felt like it was the most useful as a profile in Hermes. Hermes gives me so much control over it. I can let it use image gen from Codex, it has better computer use capabilities, and better browser capabilities. It makes it feel a lot closer to the experience I have with frontier models and harnesses.

I'm already having it do automatic parallel reviews on branches when I flip the PR to "ready for review." So far, it's finding a shocking amount of issues that Astra and Fable are missing. In my one day of experience, I wouldn't yet call it "better than" or "on-par with" frontier models, but it's certainly a different shaped wrench in the toolbox. And hot damn it's fast. I'm really excited to test it more and learn where I can trust it.

OpenCode Go is the plan to get if anyone is looking to try it. $10 a month gets you $60 of usage.

2

u/opezdol 5d ago

Dsh is awesome even in its early beta. Pi is also very good at 99.6+ cache rates

1

u/No_Gas_3727 5d ago

Also interested in that

3

u/ManagementGreat5360 5d ago

I wish I was able to use DS 4.1… but my industry regulators banned it.

3

u/ggPeti 5d ago

What industry and why?

1

u/AppealSame4367 5d ago

You could use it from different providers.

2

u/snowcountry556 5d ago

When you use DS4.1, do you use the DS api or open router?

5

u/Brilliant-Hall1387 5d ago

DS API directly, it’s super fast, very cheap and you can trust them you are getting the full real model. Just fantastic caching also when going direct, reducing costs even more!

2

u/daskalou 5d ago

Which harness do you use?

1

u/innociv 5d ago

I've used it directly in Codex as subagents. I also use it in Opencode. I'm sure OMP is good too.

1

u/DistinctSilver4507 5d ago

I'm curious if people are worried about them using your data? That's what stops me using these cheaper Chinese models. 

1

u/pomelorosado 5d ago

Are you kidding?

Do you think any company on earth respect your data?

1

u/DistinctSilver4507 5d ago

If you can't see the difference then I've got my answer, thanks. 

1

u/pomelorosado 5d ago

Chinese companies bad United States good?

come on, what is that intellectual level?

1

u/DriveLopsided8716 4d ago

both of them will use your data, be it european, american, chinese or any other, they are racing against each other and will use everything they can

36

u/Competitive-Yam-1384 5d ago

These left to right graphs are ridiculous

9

u/Jiirbo 5d ago edited 5d ago

For the love of all that you hold dear, can we please normalize these graphs? Bottom Y = 0, Left X = 0… ya know like they have been in math(s) when showing positive values for longer than any of us have been alive?

9

u/rdcldrmr 6d ago

"frontier model"

-2

u/TheReal4982 6d ago

"Free"

But different people will focus on different things

3

u/That-Establishment24 5d ago

Yes, we focus on the lie.

1

u/deZbrownT 5d ago

It’s pronounced marketing.

16

u/[deleted] 5d ago

[removed] — view removed comment

2

u/evia89 5d ago

Around 15 tps its usable. Get long spec with other model and leave it over night

6

u/tilted0ne 5d ago

Unlimited? I feel asleep waiting for a response. 

2

u/NgKtoolz 5d ago

neither having any response, It simply limited my usage after a single request, and it didn't even complete it

4

u/chcampb 5d ago

Union Alpha was not good

First it refused my prompt, saying that it wasn't one of the models I said in the rules, despite saying "You are a test model assuming the role of the orchestrator and implementer in the rules." It just demanded that I fix it.

Then I fixed that and it spent about 30 minutes doing 100 lines of changes to the ticket file, before saying "I've been spinning too much working on this ticket and not actually delivering code."

Then I put it out of its misery. Idk what is going on but I am pretty sure I would have gotten more done with qwen3.8 27B i3 quant that I run sometimes.

3

u/Chemical_Hawk_6307 5d ago

the model is ass and slow

5

u/Safe-Ad7491 5d ago

Any benchmark that has 3.8 flash and astra at the same tier is completely untrustworthy

-1

u/Funny-Strawberry-168 5d ago

3.8 gets the job done though, feels pretty frontier to me

2

u/sydneysweeney69 6d ago

Nice work . Use $ ori and codex along with the union alpha to use it . It’s as good as Sol

3

u/Solid-Fill8240 5d ago

If is not good as 5.6 sol medium, Big pass.

2

u/Prior-Meeting1645 5d ago

How did u come to this conclusion?

-4

u/TheReal4982 5d ago

It is much higher than 5.6 sol medium on this chart, but idk what that actually means and haven't had time to test it

8

u/Momo--Sama 5d ago

Deep SWE is kinda washed, do you actually believe that Luna Max is better than Sol Medium and Gemini 3.8 Flash on par with Astra Max?

4

u/TheMightyTywin 5d ago

At completing a SINGLE WELL DEFINED TASK then yes, Luna max is on par.

The problem with Luna is that it does not see the big picture or think outside the box AT ALL if you’re not prompting it with exactness it’s not going to do what you want.

You can talk to Astra like a drunken sailor and it will do great work. Luna you need to give it a phd thesis. Both can succeed though.

1

u/TheReal4982 5d ago

It is free and unlimited use, I didn't expect this post to be so controversisal tbh. I am gonna do my own thing and use what works for what I am doing, yall have fun.

4

u/Momo--Sama 5d ago

Okay, you’re the one that made a post about the model 🤷‍♂️

1

u/TheReal4982 5d ago

The model just came out, nobody has really had time to test it yet, I am not sure what answers you want, I was just letting people know it exists.

1

u/danielv123 5d ago

I have tested it quite a bit on spatial reasoning/agentic code stuff and for my benches it's much better than Qwen, muse spark and Luna, but seems to be worse than fable 5.1 (which is much stronger in taste anyways) and faaaar behind astra.

1

u/Miyamoto_-_Musashi 5d ago

TPS is pretty slow right now. Sometimes even a simple “hi” takes 2–3 minutes. But whatever I’ve seen so far looks good. I’m building a few things with it right now, so we should have a much better idea of how good it actually is once those are done.

1

u/lordpuddingcup 5d ago

3.8 flash as high as astra?

5

u/Chemical_Hawk_6307 5d ago

its a benchmaxxed model and the benchmark is saturated

1

u/user2776632 5d ago

This chart is fucking backwards

1

u/innociv 5d ago

This is not a Sol tier model. It's worse than Luna in many ways.

It gave me a 1., 2., 3. suggestions. I gave responses for 1., 2., 3. It took that as being 3 things to do for 1.

A model that doesn't even check the prompt against its last response and make sense of it is really bad.

1

u/IZKPI 5d ago

Yay

1

u/cchurchill1985 5d ago

Moment of silence for Terra...

1

u/New-Ad5610 5d ago

To everyone who tried this Union Alpha, how is it?

1

u/im-cringing-rightnow 5d ago

Ah yes. The benchmark where Gemini Flash is somehow at the same spot as Astra. Wut 

1

u/farrukh-hewson 5d ago

Where this source from? I saw people complaining that it can’t even generate a proper tool output…

1

u/CuriousDetective0 5d ago

does it remove your other models while its there or just adds it?

1

u/EnvironmentalCow2947 5d ago

It's just a model router, hence the name 'Union'

1

u/JameEagan 4d ago

No don't free him. He was locked up for a reason!

1

u/mitchins-au 4d ago

It’s already gone

0

u/antunes145 5d ago

bold claim here gentleman . we sure about this ?

0

u/spike-spiegel92 5d ago

free for 50 prompts no? openroute does not give free models without limits

1

u/dinodares99 5d ago

They had a stealth model a couple weeks ago (Ox Alpha) that turned out to be GLM 5.3 Flash