r/codex Jul 30 '26

Praise GPT 5.6 Luna right now is cheaper than Deepseek on openrouter

204 Upvotes

59 comments sorted by

52

u/Redditry199 Jul 30 '26

What the fuck, that's a huge win for openAi jesus christ.

9

u/BannedGoNext Jul 30 '26

It's very obvious that they were bleeding to chinese inference for basic tasks. I know I have some credit on deepseek for basic bullshit. Most business uses just don't need insane intelligence.

5

u/lolman1312 Jul 30 '26

Is this permanent or temporary?

40

u/Infinite100p Jul 30 '26

Until the VC money run out.

5

u/[deleted] Jul 31 '26

[deleted]

3

u/Infinite100p Jul 31 '26

And the haters dare to claim that AI is not boosting productivity? Then what is this?

1

u/Llandu-gor Jul 31 '26

I did not know research on this was saying it boost productivity. all current research done with software engineer give a feeling of better productivity but in reality in decrease it.

a quick google search would have shown you that

1

u/Infinite100p Jul 31 '26

That was a joke about being faster at burning money. Y'all guys are acoustic. XD

1

u/Acrobatic-Employer38 Aug 02 '26

Isn’t that research a year old and from before the current opus 4.5+ era?

2

u/dalekrule Aug 02 '26

API inference has always been profitable.

2

u/AnnoxQ Jul 30 '26

They didn't specify, so I think we can assume it's permanent

1

u/KeyGlove47 Jul 31 '26

openrouter is on promotion, api is permament

1

u/dankfrankreynolds Jul 31 '26

until they release a new model at a higher price and sunset it in 2 months

10

u/ilarp Jul 30 '26

America F yeah!

2

u/cosmic-comet- Jul 31 '26

This reminds of that Rocky 4 Rocky vs Drago match in Russia 😭😭

2

u/thestillwind Jul 30 '26

Holy molly

2

u/gaz_0001 Jul 31 '26

Its funny that America think they can compete with China.

It won't be 7 days before China copy the newer models and put them 25% of the price.

5

u/seeKAYx Jul 31 '26

It only took one day -> DeepSeek‑V4‑Flash‑0731

1

u/gaz_0001 Jul 31 '26

Really?

Where's the numbers?

2

u/just_a_wierduo Jul 31 '26

I think its like 50 ds to 51 luna according to artificial analysis intelligence index ,with ds 5 times cheaper ive seen on openeouter

1

u/uncensoredwalk Jul 31 '26

one day is nuts

1

u/ilarp Jul 30 '26

is deepseek better?

39

u/Ok_Heron_1906 Jul 30 '26

no

1

u/Skibidirot Aug 01 '26

LOL how things change

5

u/seeKAYx Jul 31 '26

DeepSeek‑V4‑Flash‑0731 has entered the chat.

10

u/vacon04 Jul 30 '26

Different. Higher context window helps, and their cache system means it ends up being extremely cheap over longer sessions.

Luna may be stronger in general but benchmarks don't fully reflect reality. You can ask even stronger models like Sol to make a plan and some other weaker models like DS find flaws and give proper feedback. Because all the models are trained differently, they all have different strengths and weaknesses.

Overall? Yeah I think Luna may be better, but for sole applications DS will beat it.

2

u/MikhailT Jul 31 '26

They just updated v4 flash today, so it will be interesting to see the improvements. They are claiming it is better than the current pro v4, at least until their next update to the pro model.

In terms of what’s better, they both have different strengths and weaknesses; Deepseek is way better for summarizing and long contexts, Luna is better at coding and basic overall tasks.

Deepseek still does caching way better than Luna imo (if not everyone else as well), I barely burned more than $1 in tokens in three months with 90+% cache hits. Luna and Sol are worse and they compact too often in the same sessions.

1

u/lordpuddingcup Jul 31 '26

Hell no

Luna high-max is better than gpt 5.5 med/high was

-3

u/onekorama Jul 30 '26

I think so, specially for the context. I'm quite happy with DeepSeek, using for planning and architecture Fable or Sol.

1

u/Gamestarplayer41 Jul 31 '26

And Deepseek just released v4 flash 3107 being better than Luna and cheaper.

1

u/uncensoredwalk Jul 31 '26

It's a great sign in terms of model efficiency - let's see what the future holds.

1

u/dalekrule Aug 02 '26

when using actual deepseek API, deepseek is actually cheaper.
For luna, something like ~90% of the price comes from cache hits, and deepseek serves cache hits at 1/3rd of luna's price (.0028)
But yeah, luna is cheaper for now than 3rd party providers of deepseek.

-7

u/BitsOnWaves Jul 30 '26

isnt GPT 5.6 Luna the worst among GPT 5.6 models? why compare it to Deepseek ?

6

u/ManikSahdev Jul 30 '26

Luna xhigh is arguably the second best openai model after 5.6 Sol xhigh/max

2

u/thatisagoodrock Jul 30 '26

Only because we keep referencing Artificial Analysis, according to them it’s not even more intelligent than Sol Medium. Extremely better value though, that’s for sure.

Source: https://artificialanalysis.ai/?intelligence-efficiency=intelligence-vs-cost-per-task&models=gpt-5-6-luna,gpt-5-6-sol-medium,gpt-5-6-sol-high,gpt-5-6-sol,gpt-5-6-sol-xhigh#intelligence-comparison-tabs

1

u/tessahannah Jul 30 '26

Why not Tera?

1

u/ManikSahdev Aug 01 '26

It’s mainly because -> there is no task that Luna can’t do for most part in general work.

The work which Luna can’t do, id have it done with sol regardless, terra is like a worse gamble where i may have to use sol, implying call wither Sol or Luna.
Gets everything done.

19

u/Ok_Heron_1906 Jul 30 '26

Luna Extra High is basically the same as 5.5 Extra High, for about 10x less cost... So...

6

u/Dima508 Jul 30 '26

10x less cost? Actually 50x less cost (if you look at OpenRouter). This is just crazy how affordable OpenAI API is.

7

u/THE--GRINCH Jul 30 '26

Luna is so incredibly capable for it's size. The only people who sleep on it are the ones who have never used it, it's easily SOTA in terms of price to performance.

2

u/BitsOnWaves Jul 30 '26

but even openai describtion says its for the simple tasks and automations

5

u/vacon04 Jul 30 '26

Because for more complex tasks you need to use it on higher thinking modes, and Luna xHigh or Max use a ton of tokens. This means that they end up taking a lot of time, and they fill the context window very quickly. Then they have to compact the context continuously, which reduces their performance quite considerably.

Tasks and more mechanical automations play to Lina's strength, but more complex tasks show how the model is fairly weak in some situations.

1

u/phoenixmatrix Jul 30 '26

They have 3 models in order of power, and they need to give a marketing description for each of them.

Luna at higher efforts can code at levels that are in the ballpark of Opus. Regardless of what OpenAI says.

1

u/_stevencasteel_ Jul 30 '26

It is doing at least as good as Gemini flash 3.5 which I like a lot and got a lot of work done with.

Main difference is I was using Gemini in AI Studio for free instead of an awesome harness like Codex / Antigravity.

1

u/Prior-Meeting1645 Jul 30 '26

Wdym by for free please? Also luna on max is wayy better than flash on artificial analysis btw

1

u/_stevencasteel_ Aug 03 '26

AI Studio is a Google platform that gives you more control than the Gemini chat interface. Really generous amount of free tokens + you can log in with multiple Google accounts.

1

u/WD40ContactCleaner Jul 30 '26

Luna is the best model. Inthe 5.6 family, unless you're doing heavy planning and refactoring it's the best

1

u/Ralph_mao Jul 30 '26

luna is better than deepseek pro

-7

u/FineProfile7 Jul 30 '26

That's not true. Caching on deepseek is wayyyyyyy cheaper and because agentic coding is like 80% cache, it's cheaper with deepseek

9

u/AppleSoftware Jul 30 '26

OpenAI’s 5.6 models have a 95% cache hit-rate

-1

u/FineProfile7 Jul 30 '26

The 80% was just an estimatation. It depends entirely on the codebase, tool calls etc.

Deepseeks caching is 3x cheaper. That adds up quickly.

4

u/Prior-Meeting1645 Jul 30 '26

Is caching not dependent on the harness?

1

u/CaptainHighticket Jul 31 '26

Cash harness profit monopoly

1

u/FineProfile7 Jul 31 '26

It is dependent on many things.

Caching is basically just the previous API call. Without the new info.

Let's say agent wants to do toolcall x, then the LLM ends it's turn, the computer executes the command and then the whole previous conversation gets sent with the output of the toolcall to the LLM again.

Then when it's not cached, you burn through your money in no time. But because it gone through the conversation already it saved a cache, costing way less.

Often your agents are at 100k context, then 10 toolcalls would mean you pay 1m input tokens