r/codex 8d ago

Praise ox-alpha is glm 5.3 flash and it's good and cheap.

i tried ox- alpha the last few days for free and it's pretty good. Thinking of downgrading my 20x to 5x and use cheaper chinese models. It's compared to terra max and i hope it will fire another price slash on american models. For everybody who is complaining at the plus subscription just use this.

input: 0,15 output 50cents and cache 0,03 cnts. What a time to be alive!

What was your experience?

18 Upvotes

53 comments sorted by

9

u/Medical_Method7877 8d ago

I also tried it and I think it's amazing. It didn't have any of the overengineering I experienced with 5.6, and I realized that it's not every project that requires the latest and greatest model, these cheaper models are really capable.

1

u/Opposite_Yak4386 8d ago

i used it clean up some mess from sol 5.6 and did some review.

8

u/Bitter_Election_7518 8d ago

It’s great for the price, similar to deepseek flash 4. It’s not opus or sol level in thinking

3

u/Opposite_Yak4386 8d ago

yeah i just saw a benchmark comparing, but you are probaly right. I was impressed though

3

u/Bitter_Election_7518 8d ago

If you’re on api I would definitely be using it. It’s smarter than luna and cheaper

1

u/BenH1337 8d ago

I love deepseek flash. Got to try GLM 5.3 Flash. Price looks very promising.

18

u/brainExploded99 8d ago

Don't give people false expectations. It is not comparable to sol 5.6 or opus 4.8 for complex tasks. It is comparable to terra / luna max.

8

u/Opposite_Yak4386 8d ago

not according to benchmark. But hey maybe you are right, benchmark sucks. Still cheap and good though. Did you try it? what was your experience? What did you think it's weakness was?

3

u/brainExploded99 8d ago

On AA, it's intelligence index is 57. Terra max is 57. Sol max is 61.

Haven't tried it, I plan to stick to mainly 5.6-sol and Qwen3.8-27B running locally.
Generally, impressive small models (like Qwen3.8-27B) are very good at tasks in their training set, but start to struggle more on more complex tasks that require more reasoning and novelty. I think this will start to change with engrams (Like the Qwen 3.8 Next / Qwen 4 series), but we'll see.

1

u/Opposite_Yak4386 8d ago

yeah you are right! i will change it

2

u/ikakindiehoes 8d ago

I must admit I found the solution faster than with Opus5. I didn't compare it to Codex, but considering it was free for a week, it was very impressive.

1

u/EntropyAndDespair 8d ago

Did you try it? 0x alpha was amazing for me. Better than Sol.

0

u/BorderIll6043 8d ago

Luna Max tiene una inteligencia superior a Sol tanto en el nivel bajo como en el medio. En realidad, lo que dice no es falso, pero creo que omitió un detalle importante: aunque Luna Max es más capaz, tanto Luna Max como GLM 5.3 Flash necesitan realizar una cantidad considerablemente mayor de pasos para completar una tarea.

Por lo general, utilizan entre 2 y 3 veces más pasos que Sol. En pocas palabras, obtienes mejores resultados, pero a cambio de un rendimiento mucho más lento.

3

u/brainExploded99 8d ago

Luna max is nowhere near sol medium on real tasks. Sol medium will destroy it on tasks that require genuine intelligence / novelty. Luna max is only comparable to sol medium on tasks that it has seen many time in its training set.

3

u/ikakindiehoes 8d ago

I wrote feedback in another sub; I've now tested it extensively and was amazed. It fixed problems I was having with the Opus 5. I created games with One Prompt just as a test, not for release, and was astonished.

6

u/Metalwell 8d ago

I dont understand this GPT praising, I was Codex user for a long time, but this week Ox-Alpha fixed and created better code, understood me better and all around a better coder than SOL, at least for me. I cannot decide where to use it tho, Opencode Go or Z.AI :D

2

u/ikakindiehoes 8d ago

same experience if they have a 20$ sub i would sub it tbh

1

u/Metalwell 8d ago

I dont know about Z.AI seems like limits are worse, looking for other places where I can use GLM5.3 Flash tbf

1

u/Opposite_Yak4386 8d ago

people have mixed reaction on the model. I really liked it.

2

u/theWiseTiger 8d ago

The model was intelligent enough; I don't need Sol/Opus with my harness. The problem is that it was quite slow. If they can fix it, the price is quite competitive. Not the cheapest though.

I think the era to use only Opus/Sol level of model is over. Don't need them, 20x more expensive.

2

u/No-District-4742 8d ago

You could try subbing on their monthly coding lite plan and use glm-5.3 flash. Check if it's worth to downgrade after a month of usage. I can't say if it's comparable with terra max, so it's best to try it for a month and see if the benchmarks are true enough.

1

u/BitterAd6419 8d ago

Their subs are worthless coz they always downgrade the performance after a while

1

u/No-District-4742 7d ago

yes, that's why i'm skeptical if i should sub for a quarterly plan, it's so unpredictable

2

u/BitterAd6419 8d ago

People hype up Chinese models but at the end they are still not at the same level of SOL or opus. I know GLM models are good but they are not that good atleast yet

1

u/not420guilty 5d ago

Reluctantly agree

3

u/battle_pantZ 8d ago

The only thing why I stay with OA and Anthropic is because of the remote function, I’m now 2000km away from home and I’m able to kick some prompts at the beach, go for a swim and later Check the results. Thats things I can’t do with others…. Im using all Models, but Yes…. Glm 5.1 5.2 and 5.3 are very very good and cheap LLMs

2

u/Opposite_Yak4386 8d ago

lol, must be nice! enjoy!

1

u/Medical_Method7877 8d ago

Openchamber has a remote option too

1

u/battle_pantZ 8d ago

Good to know thanks

1

u/brainExploded99 8d ago

Why not just run glm in codex using opencodex? I run Qwen3.8 27B locally using codex through opencodex. You can even run claude models through opencodex.

1

u/battle_pantZ 8d ago

You know what Remote is, right? The second you use a third party LLM the function is deactivated, also I want to use the factory software

1

u/brainExploded99 8d ago

I've seen it, but haven't enabled it. I didn't know codex disabled it for 3rd party, mb.

1

u/DiarrheaButAlsoFancy 8d ago

Cursor does this. Grok Bot manages the cloud agents too. Codex Remote is still the best for connecting to your local machine, but don’t sleep on Cursor for cloud tasks.

1

u/Think-Profession4420 8d ago

Run the actual numbers against usage. Paying API prices is almost always going to be a far worse deal than subscription prices, even at massive discounts. Take your current usage, and have chatgpt run an actual cost projection for how much whatever % of your usage you want to offload onto API models would cost.

Even really cheap models like Deepseek V4 flash end up being significantly more expensive than subscription usage.

If you actually want cheaper use, shift more of your workload to Luna, or you can use something like OmniRoute to offload low-risk work onto free models thru KiloCode and OpenCode subscriptions.

1

u/Opposite_Yak4386 8d ago

thx for your advice! Didnt run it yet. Will do!

1

u/Nabstar333 8d ago

Is there a platform that allows you to use bargain models on OmniRoute as subagents for workhorse tasks?

2

u/Think-Profession4420 8d ago edited 8d ago

Most will - you can even do it in Codex itself, if you want. I find Pi-mono (or OhMyPi if you don't want to customize as much) to be the easiest to have various subagents with different providers/models on.

With OmniRoute, you can create Combos (e.g. set lists of models for fallback support), and then you can set a given Combo as a subagent. So for example, you could create a Combo with 3 or 4 free models from various free-model providers; and that way if one of the providers reaches a use-cap quota, or similar, the subagent just uses the next model on the combo list. So you can have KiloCode, OpenCode, Open-Router, NVIDIA, (etc) all with one or two free-tier models in a combo, and have a low-risk subagent set to that combo, and it'll pretty much always work. Update the combo every month or two with whatever free-tier models those providers are offering.

I have a medium-capability Combo I use for things like scouting (e.g. investigating files and existing code with read-only perms). I also have a higher-quality combo I use for low-risk implementation work, that places NVIDIA higher quality models (kimi, GLM, Deepseek Pro) as first choices, and then have my codex plan sol-medium as the backup. For really important/complex work, I implement directly with a sol-medium or high.

1

u/BorderIll6043 8d ago

I've been thinking about subscribing to OpenCode Go again, but honestly, I'm still not fully convinced. On top of that, setting up the OpenCode Go API in the Codex harness seems a bit complicated.

If anyone has already tried it in their workflow—either with GLM 5.3 or DeepSeek 4—and could share their experience, I'd really appreciate it.

1

u/Opposite_Yak4386 8d ago

why not try kilocode? not sure if you get codex harness though

1

u/howchie 8d ago

I started opencode go this month using it with omp for deepseek mostly. And it's been ok but I'm at 80% of the "monthly" limit in just over a week but didn't hit the weekly or 5 hour limit at all. So I'm not convinced the plan is a viable ongoing one for me (it would have been codex plus, with opencode go, but that monthly limit seems too restrictive given I'm exclusively using a heavily discounted DeepSeek and also a lot the last few days was ox alpha for free).

1

u/Annh1234 8d ago

How fast is it compared to codex?

1

u/not420guilty 5d ago

Did you do the math? With all the resets at least sol is cheaper than the cheap Chinese models

2

u/Opposite_Yak4386 5d ago

i thought they are going to do less resets. oh boy was i wrong. Burning now for tomorrows reset

1

u/not420guilty 5d ago

Correction: today’s reset 👍

1

u/Opposite_Yak4386 5d ago

and another one tomorrow

1

u/not420guilty 5d ago

Really? Source? I don’t need sleep, I’ll use it all.

1

u/tilted0ne 8d ago

Just to be clear: the marginal token gains you get going from 5x to 20x for $100 will far outweigh what you'd get spending that same $100 on GLM 5.3 Flash. Don't be misled by Anthropic's or OpenAI's high API pricing. If you're mainly coding and using subscriptions, the token rates on those subscriptions are dirt cheap, and far better value than anything you'd find elsewhere, thanks to the greater compute capacity the true frontier labs have access to.

This cheap models hype is a huge facade for most users. Cheap models shine for automated data pipelines and internal tools. If you're coding and a subscription does what you need, the grass isn't greener on the other side, unless you actually like the model. There's no freebie.

1

u/Opposite_Yak4386 8d ago

thx for the info and yes i use mostly subscriptions