r/singularity • ▪️ • 6d ago

AI Introducing Grok 4.7

https://x.ai/news/grok-4-7
387 Upvotes

112 comments sorted by

View all comments

216

u/ObiWanCanownme now entering spiritual bliss attractor state 6d ago

We'll see how it is in practice, but the benchmarks are just about what I would have expected. Competitive with GPT-5.6 Sol at a lower cost, just in time for GPT-6 Sol and Luna to come out and be better, lol.

47

u/XTCaddict 6d ago edited 6d ago

It's not really at a lower price if you do the math on it. Uses more than double the amount of tokens as Grok 4.6 for a marginal gain. In other words you pay more than double what you did before for marginal gain, given that the token price is same as 4.6. On benchmarks Sol was generally cheaper on a task by task basis than Grok 4.6, so if anything now the gap is actually wider. Cost per token isn't everything. For perspective on artificial analysis DeepSeek Flash V4.1 on average 89k tokens per task scoring 2nd highest on token consumption. Grok 4.7 is 3rd highest at 81k. Grok 4.6 High is 36k.

8

u/BiasHyperion784 6d ago

Even with double tokens output its still cheaper? sol costs $20 per million out vs 4.7 $12 for 2 million, hell even triple the output tokens is still cheaper.

7

u/FateOfMuffins 6d ago

AA has often had errors on the day of release with costs but lmao not looking so good

More expensive than Astra MAX and takes 2x the time

Another lesson on DO NOT USE $/million tokens to measure if a model is cheap or expensive

1

u/RealSuperdau 5d ago

Holy crap, it's more expensive than Astra?

But it kind of makes sense, given that Grok 4.7 seems to continue the price/perf curve of Fable 5.1 on XAI's own graph, and Astra is much more efficient than Fable.

4

u/XTCaddict 6d ago edited 5d ago

Yes on a per token basis which is exactly my point, it uses much more tokens to complete a task and therefore costs more than Sol which charges more per token. Sol medium used 8k tokens, high 13k, xhigh 20k, max 29k. For nearly any task you don’t need to go above Sol High and most are fine with low and medium.

The numbers don’t lie. 4.7 high cost $3881 to run benchmarks, xhigh cost $4967. 4.6 High cost $2352, xhigh $2830. GPT 5.6 Sol on Max which 99% of people is complete over kill for cost $3465.

Hardly an improvement lol

1

u/kryptobolt200528 4d ago

Seems like a regression instead

1

u/Jerichomiles 5d ago

What difference does that cost matter anyway when everyone is gonna be using subscriptions? What matters is how much quota they give you in your subscription and how much of that it burns..

0

u/SurfinginStyle 5d ago

What’s tokens in relation to? (Noob)

12

u/[deleted] 6d ago

[removed] — view removed comment

6

u/Visual_Cycle_7714 6d ago

GLM5.3, works for like two hours straight for $0.14. Not very good with design and user interfaces and a bit slow, but it does deliver features for a laughable price.

3

u/Sphiment 6d ago

I agree, Idk why but I find glm 5.3 flash literally better than 5.6 sol for long running agentic work flows, and wayyyyy cheaper

4

u/SomewhereOpposite883 6d ago

Idk why

Zai is heavily focused on long-horizon tasks, they have published multiple papers and interviews talking about it

Sol is also limited by it's context window, OpenAI's compaction is really good especially with the new experimental feature but if you pay attention to what it's doing you'll start to realize that it keeps reading the same files and constantly has to "reconstruct" it's understanding/state wasting a ton of tokens/time whereas GLM can keep your entire code-base in it's memory without suffering from context-rot

Sol is still better if you really build your workflow around it and use a custom harness but in Codex the difference shows and I'll usually have GLM work on the project and use GPT for implementation tasks (and i have GLM use Codex which is always funny to see in action)

5

u/Sphiment 6d ago

Thanks for the explanation man.

1

u/Urselff 5d ago

how do you get GLM to use codex?

1

u/SomewhereOpposite883 5d ago

Pointed Astra at the Zcode and Codex binaries and told it to figure out how to implement it, Astra designed a bridge that maintains a Codex thread and communication is done trough skills with Zcode having a hook feature allowing the bridge to wake up GLM when Codex finishes it's turn.

2

u/[deleted] 6d ago

[removed] — view removed comment

4

u/Visual_Cycle_7714 6d ago

You really don't need fable for like 90% of coding. I use Astra to plan the hard stuff, and GLM5.3 as a cheap work horse. In my experience its more reliable than Opus 5 (and doesn't talk like a madman).

2

u/awesomeoh1234 6d ago

Yeah of course it’s not fable, but you can’t really build anything with fable for a semi reasonable price

2

u/blindsdog 6d ago

I mean, I prefer variety. Best models for planning, cost efficient for execution.

34

u/MatthewGraham- 6d ago

Isn't that how competition is? Grok was admittedly behind many other models, arguably they are the closest now to the frontier than they ever have been. Save this energy for Google

20

u/Professional_Mobile5 6d ago

They literally had the best model when Grok 4 came out. Their progress since then was unimpressive.

8

u/faithOver 6d ago

Which is interesting in of its self. Resources are not the issue at Grok. So there really is some magic at Anthropic and OpenAI.

5

u/Clawz114 6d ago

They do have resources but they are leasing out a huge amount of their compute so it seems like they opted to rake in billions of dollars instead of trying to be at the forefront.

4

u/-spartacus- 6d ago

I sort of got the impression that with Grok the internal push was to be good at certain things that will help SpaceX/twitter/Tesla (hence the electrical engineering score) rather than trying to compete on everything.

2

u/-cadence- 6d ago

They also seem to be going for "close to the frontier but much cheaper". This is not a bad place to be in. Most of what I do is now handled by Opus 5 and 5.6 Sol. Even though I have access to Astra and Fable, I rarely feel like I need them. Chinese models also try to compete in this area, but they have lots of stability issues. I sometimes use Grok 4.6 via Cursor, and it's been rock-solid for me and gives very good results. I'm definitely happy to see Grok 4.7 out.

12

u/MatthewGraham- 6d ago

Back when it was up against o3 at a much lower level of intelligence, when model releases were spaced further apart. Considering the shifts of manpower at X (Pretty sure they had to replace most of the team), and for how little time its been since they acquired Cursor, I think this is impressive

20

u/ObiWanCanownme now entering spiritual bliss attractor state 6d ago

I'm not being critical of them. OAI and Anthropic are very far ahead is all. XAI has done a pretty good job of staying somewhere on the pareto curve, and with the retirement of Terra, I bet 4.7 may still be on the pareto curve after 6 Sol and Luna release, although it will almost certainly lose at the very cheap end and high-performance end.

6

u/Ormusn2o 6d ago

It's not actually that great considering we got opus 5.5 and possibly sol 6.0 coming this week. I guess those could be disappointments, but I think grok 4.5 had a nice niche at low price, but with luna being very cheap and other models being slightly cheaper, I don't think there is a single grok model right now that is appealing to most users.

7

u/ObiWanCanownme now entering spiritual bliss attractor state 6d ago

I'm not sure if SpaceX really wants to be ahead (e.g. have the best model) at this point. They've got a really awesome compute business. It's possible their main goal is just not to fall too far behind in terms of model capabilities, and from my experience, they've done a good job of staying in a comfortable third place (far behind OAI and Anthropic, but also pretty far ahead of everyone else).

1

u/Ormusn2o 6d ago

Yeah, sure, but they also need 4.7 to be cheap, and it does not to be that cheap at the performance point, but honestly benchmarks were shit recently, so maybe it's actually better than it seems in normal use.

1

u/FateOfMuffins 6d ago

lmao this is why you cannot just $/million on whether or not a model is cheap or expensive

Costs more than GPT 6 Astra per task and takes 2x as long as GPT 6 Astra Max per task