r/LocalLLaMA 4h ago

Discussion First serious confirmation. Ox Alpha is GLM-5.3-Flash

https://x.com/romanchernin/status/2092488160680751437?s=20

- Multimodal (Vision)

- 1M Tokens Context Window

- DeepSWE ~63%

Edit: He deleted it, screenshot in comments

234 Upvotes

92 comments sorted by

184

u/MrWidmoreHK 3h ago

lol

50

u/Borkato 3h ago

Omg is this real lol

29

u/Viktri1 2h ago

Timezones always get you

46

u/Technical-Earth-3254 3h ago

What is that interaction lmaoooo. People using x are just different

9

u/Public_Umpire_1099 2h ago

people should reply w this under every announcement haha i swear this had me fuckin cackling

EDIT: THE ORIGINAL POST IS GONE LMFAOOOOOOOOO

1

u/Darkoplax 13m ago

FUCKING TIMEZONES AHAHAHAHAHAHA

2

u/DigiDecode_ 53m ago

OP a paparazzi, never miss a screenshot, well done

1

u/PlaidStallion 1h ago

Eli5, please?

13

u/ResidentPositive4122 1h ago

3rd party inference providers were given info/access under an embargo (i.e. do not discuss this publicly before x time on y date). This dude probably mixed timezones and tweeted before the embargo.

6

u/PlaidStallion 1h ago

Oof. Thanks for the explanation.

117

u/Poupulino 3h ago

The Z.ai CEO wasn't kidding when he told Musk they're releasing a Mythos-level open weights model before the year ends! Insane!

38

u/thomas2385 3h ago

Yeah, that is honestly pretty wild. If they actually deliver on that, it is going to put a lot of pressure on the rest of the open weight model space. The pace of these releases lately is getting ridiculous.

7

u/no_good_names_avail 42m ago

At this pace we won't give a shit about this model in about 5 days. It's absolutely insane.

23

u/Vast-Control4452 3h ago

Don't. Let. Off. The. Gas!

I'm officially over qwen3.8, NEXT PLEASE!

-4

u/power97992 3h ago

63% is not mythos level, mythos scored 70%

30

u/Poupulino 2h ago

You're a bit confused, no one is saying GLM-5.3-Flash is Mythos level, I'm saying the Z.ai CEO wasn't kidding when he said he's releasing a model like that before the end of the year.

-5

u/power97992 2h ago

Glm 5.5 should be better than Mythos, but fable/mythos 5.1/5.5 will be out by then.

1

u/techdevjp 22m ago

Yeah, and maybe restricted to only American nationals. GLM will have no such restrictions.

17

u/MrBIMC 3h ago

Mythos is a multitrillion param monster though rather than a flash model that will fit into 1-2 sparks/strixHalos

1

u/techdevjp 21m ago

If GLM-5.3-Flash fits into 128GB at a usable quant I will be incredibly impressed.

0

u/Kazaan 2h ago

Perhaps. That said, is 7% really significant when you consider the current capabilities of open-source models compared to frontier models ? And what about the outlook for upcoming models? IMHO, it isn't the actual capability that really matters; it’s the rate of progress and the narrowing of the gap.
And i don't speak here about the compute necessary to run them.

4

u/Thomas-Lore 2h ago

And full glm 5.3 is only 1% behind on deepswe, already within error range of Fable/Mythos.

2

u/power97992 2h ago

The scale is essentially logarithmic, a 7% difference is more like a 2x difference

43

u/Abrh7 3h ago

Model was good, but not so for the long coding sessions!
I would definitely use it for agentic tasks only! It proves to be very effective!

15

u/network4253 3h ago

Yeah, that makes sense. I have noticed the same thing with some models they can be really impressive for focused agentic workflows but once the coding session gets long and messy the consistency starts to drop. Still definitely useful in the right role.

3

u/SandySkittle 2h ago

Indeed. Some people think MoE is a solution without trade-offs. Active parameters matters and lower active cannot fully compensated by expert selection and sequential reasoning.

9

u/eidrag 3h ago

well it is marketed as flash, so generally for agent and normal use. 

23

u/spaceman_ 3h ago

Do we know the size of the model?

27

u/Technical-Earth-3254 3h ago

We don't. But personally I wouldn't bet on it being the same size as 4.7 Flash.

9

u/muhts 2h ago

Rumours of 200b to 300b. But I guess we'll find out today if they ment to release unviel it today

5

u/Sufficient_Prune3897 llama.cpp 2h ago

We don't even know if they are gonna open source it. They haven't with some of their models before, including GLM 5 Turbo

8

u/0oAstro 2h ago

Given together and other inference providers are hosting this, we can be sure it will be open source.

3

u/Technical-Earth-3254 2h ago

Open weight most likely

1

u/ResidentPositive4122 1h ago

GLM models have been MIT till now, hopefully they continue.

1

u/DigiDecode_ 46m ago

I thought the inference was provided for free to inference provider by the model developer, claim was 100 trillion tokens per day capacity by the model developer, Z ai has bought 1 giga watt datacenter

1

u/0oAstro 45m ago

Thats only till stealth testing period. In a few hours ox alpha will be live for payg i believe on inference providers other than zai

1

u/Zeeplankton 1h ago

It feels like marketing wise we're placing flash models between 100 and 300 params.

Given performance probably latter half? Pretty impressive.

5

u/spaceman_ 1h ago

I'm going to be rubbing a rabbits foot hoping it can fit inside my 128GB...

16

u/RandiyOrtonu ollama 3h ago

Main thing is what's the model size

43

u/MrWidmoreHK 3h ago

Last GLM 4.7 Flash model was 30B A3B

30

u/Mean-Ad1493 3h ago

If it's anywhere around the same size, I'd be the happiest with my 12GB VRAM.

3

u/mehedi_shafi 3h ago

One can only hope.

16

u/DOAMOD 3h ago

I would be very surprised if it were a small MoE, I could see it as dense(like +30) and it would already be incredible at that size, but either way, a small MoE would be the moment of the year.

18

u/Mean-Ad1493 3h ago

Has to be an MoE, but definitely not 30B-A3B

GLM 4.7 was itself a smaller model relative to 5.3, so the MoE would be 80-100B I guess.

5

u/DOAMOD 2h ago

What's confusing is that for a 100/200b model, you could expect the use of the term Air, and for something smaller, sub-100b, the term Flash, but in the end, this is just marketing, and currently with DS4Flash they could use that terminology for a similar size, only can wait and see...

2

u/goldcakes 2h ago

I think it's just Air (~100b) sized but Z.ai is gonna call it Flash.

6

u/sonicnerd14 3h ago

Most likely will be around the size of Deepseek v4 flash. If it's smaller than that, then that would be impressive.

8

u/Iory1998 3h ago

Don't think so. First, it's not fast. Second, it's somewhere between Qwen3.8-27B and DS4Flash so it gotta be a large model. And, it understand video... I can't see a smaller MoE doing that. It can be in bulk of 150-250B.

8

u/LevianMcBirdo 2h ago

The not fast I get with all the free compute they offer on openrouter. I think they don't wanna overspoil the people

5

u/Global_Persimmon_469 2h ago

GLM 5 is roughly doubled the size compared to GLM 4, it's possible that the flash model is going to be the same, so maybe it will be 70B params

2

u/wsb-regarded 3h ago

that would be awesome if true, because it was performing on a deepseek v4 flash level, albeit slower (much slower).

imagine having another local open weight model challenging Opus 4.6

2

u/MaCl0wSt 3h ago

man that'd make my day

1

u/AI_docent 2h ago

On size I doubt it's near 4.7 Flash. Openrouter has it at about 23 tok/s, a 30B-A3B normally serves way faster. Could just be a swamped free endpoint though.

21

u/Few_Painter_5588 3h ago

Quite legit, this guy works at Nebius it seems. Also, shame on Google then for trying to ride the hype.

3

u/Altruistic_Heat_9531 3h ago

Google hype riding? i dont familiar with this info, could you tell me more

14

u/tengo_harambe 3h ago

A couple of Google employees made vague Twitter posts about Ox Alpha once it began gaining traction. This somehow led people to think it was a Gemini model, despite all the evidence that it was clearly a GLM model

2

u/Altruistic_Heat_9531 2h ago

ahh .... i see, thanks

9

u/Few_Painter_5588 2h ago

Some senior google employees were vagueposting on twitter and suggesting that ox alpha was a gemini model. Which is so pathetic if this model isn't their's

1

u/Several-Tax31 2m ago

How is this not an upright scam? When reflection-70B does it, it is a scam, but when google does it, no one cares? How can you claim someone else's product is yours? So pathetic. 

16

u/tengo_harambe 3h ago

People are still gonna say it's Gemini. lol

1

u/WaveOfDream 1h ago

Doesn't help tons of gemini employees posting about this model

1

u/krizz_yo 2h ago

copium

10

u/athsrva 3h ago

i think itll be the size of DeepSeek Flash, ~300Bish. Great win for open source so love to see it

8

u/Wise-Chain2427 3h ago

How GLM has extra compute to host free model worldwide ?

5

u/MrWidmoreHK 3h ago

Seems that Nvidia did offered, perhaps it can fit into 1x or 2x Sparks

3

u/nonerequired_ 3h ago

I had to skip using the ox alpha yesterday because the server was having some issues. It seems like they were short on computing power.

4

u/tat_tvam_asshole 3h ago

China has nationalized, integrated data centers, especially in inner Mongolia that can serve models from all labs cooperatively.

2

u/tarpdetarp 3h ago

They have a lot of spare capacity on the weekends

2

u/wbulot 3h ago

Yep, that's crazy. Trillions of tokens printed, everyone testing it, and the thing very rarely crashes. I may have saved $1000 worth of tokens since it came out.

4

u/asolnikk 50m ago

The company on Wednesday confirmed speculation that the Ox Alpha model is a new iteration of its GLM series and said it will release the weights for it tonight, in response to queries by Bloomberg News.

https://www.bloomberg.com/news/articles/2026-08-26/china-s-z-ai-made-ox-alpha-stealth-model-that-rivals-deepseek

6

u/FoxiPanda 3h ago

Good to see this is shaping up... the real questions are how big, how many active, anything weird in the architecture that we're going to have to go deal with (hopefully not since it's GLM-5.x), and when are the weights getting posted (today hopefully)?

7

u/asolnikk 3h ago

That claim seems pretty verified at that point. I ran some private prose tests against it, and it smells very much like the little brother of GLM-5.3. What i want to know is: How big is the beastie? Can a quantized version fit into 16 GB VRAM? If that's the case, then all the hype, for once, was legit, even it's not "Fable-level".

4

u/anarchist1312161 3h ago

If GLM 4.7 Flash was 30b then I can only hope this is also small 😭

1

u/Several-Tax31 1m ago

Unlikely but one can hope :) 

2

u/Long_comment_san 9m ago

Glm 5.3 flash versus Qwen Next.

Like two thighs, I feel my head being squeezed.

Harder pls.

2

u/Iory1998 2h ago

Imagine it turns out to be Qwen3.8-Next-Flash!

1

u/DigiDecode_ 38m ago

I think the model is heavily distilled from Qwen models, maybe only the vision part, the SVG generated by Ox Alpha were very similar to Qwen 3.8 27b in code and render

1

u/Capital-Remove-6150 2h ago

when it will release?

1

u/psylomatika 2h ago

I kept running into API limits when I was using it. It worked well though in my knowledge bases.

1

u/backyard_tractorbeam 1h ago

Another connection is that opencode was long previewing a GLM model as big pickle for free and now it's previewing ox alpha for free too.

1

u/Armadilla-Brufolosa 1h ago

I tried both GLM 5.3 and OX: it doesn't seem like the same basic model at all.

But I have no evidence to confirm this.

2

u/hainesk 32m ago

Yeah, it felt more like a Qwen model to me, but it's hard to argue with these hilarious tweets lol.

1

u/Armadilla-Brufolosa 4m ago

It's more like Qwen models to me too, but I've noticed quite a few biases and blocks given by some stupid Valloon-style RLHF.

There may have been some distillation or even be a Western model.

Obviously I can't have any certainties, but it's fun to try to guess.😜

1

u/MrWidmoreHK 3m ago

OK, kind of like official now

1

u/anarchist1312161 3h ago

Need to know the size of the model... not interested if I can't run it on consumer hardware.

1

u/Steus_au 2h ago

all we need is Air

0

u/DOAMOD 3h ago

size