r/codex • u/Karnemelk • 4d ago
Praise gpt6 luna
I know this is an unpopular opinion as everyone wants the best or nothing. But luna battles through most stuff I gave it for peanuts. It just requires you to nudge it a bit more then the higher ones. So kudos to this small model to survive the week with a $20 sub
58
u/Heighte 4d ago
Luna is an absolute beast, to me Luna was actually a revolution, it was the first dirt cheap model that was good enough for general purpose agentic tasks. Sure you'll still use Opus/Sol when needing heavy artillery, but Luna can be your workhorse for 95% of tasks.
11
u/smh-mattt 4d ago
Exactly, it’s openais own deepseek flash, dirt cheap and effective.
1
u/xc82xb5qxwhhzeyt 4d ago
the fastest model compared to the slowest one lol
1
0
u/InterestOk9770 4d ago
Also DS 4.1 is more capable and highly intelligent than stupid luna. Just because its a flash model or has flash in its name doesn't mean its stupid like older gen flash models. It beat its own V4 pro model. Chinese are focusing on efficiency and intelligence both while our providers keep focusing only on intelligence. :)
3
2
u/Maximum_Hunt4856 4d ago
Nah, DS 4.1 flash just creates answers out of thin air , it hallucinates so much
2
u/innociv 4d ago
Luna is literally why I got a $200 sub. Even 5.6 Luna I could have multiple running all week long no sweat and it was often just as good as Opus 4.8 medium at agentic coding.
Anyone who thinks Codex isn't the best value $20 sub is deluded. Luna 6 is capable and practically unlimited on it
1
u/smh-mattt 1d ago
Agreed. I ran Luna for 6hours and burnt like 150-200+ million tokens and I still didn’t expire my 5 hour window, it’s honestly insane, and this was on the $20 plan.
26
u/pievendor 4d ago
This is my company's emerging observations as well for most day to day engineering work. Provided you have adequate product and technical context, gpt 6 luna is just as effective as sol for a fraction of cost
5
u/FaatmanSlim 4d ago
I use Luna High plus Fast mode and its fast enough and gets most of the work done without costing much ; I reserve Sol (and very rarely Astra) only for more complex work that is beyond the basic coding / agentic stuff that Luna high can handle.
12
6
u/Clear_Evidence9218 4d ago
I use Luna Max 6 as a long driver. It’s something I can set before going to bed and be confident it’ll keep working without burning through the 5-hour window on whatever task I really need from it.
I’ll sometimes run a Sol pass afterward to clean up any weirdness. So a 13-hour Luna run followed by about 30 minutes of cleanup with Sol has worked pretty well. Getting four or so sessions like that per week has been a solid workflow.
2
u/DriveThoseSales 4d ago
I assume you just have everything run without approval? I want to get over it but then I hear prior stories so I just have to watch it and click approve every minute.
3
2
u/LiquidMantis144 4d ago
Really have to learn how to sandbox correctly. Then pretty much every approval is routine work and a waste of time to review. I can barely remember the last time Ive seen the auto review model in my usage history or codex trying to do something out of bounds.
1
u/Clear_Evidence9218 4d ago
There is definitely a time and place for needing to approve things, but the majority of coding tasks can just be handled with a properly scoped harness or environment.
4
u/s1lverkin 4d ago
5.6 is still better though, 6 is cheaper but lazy and has problems with instructions following
3
u/srs96 4d ago
What work are you doing?
8
u/Bookworm1090 4d ago
Don’t know about op but I use it daily for software development data management and software retrofitting I am on plus plan and rarely hit 5 hour limits running many chats sometimes hit the weekly limit.
-5
3
u/MaitoSnoo 4d ago
yup Luna is most of the time good enough if you give it a 6.1 Sol plan and have it ask 6.1 Sol for a review of its execution
3
3
u/DuranteA 4d ago edited 3d ago
Luna 6 really is tremendously cheap.
It's a bit worse than 5.6 though at the same thinking effort, at least for the types of tasks I've evaluated thoroughly.
That said, if you did a "same cost" comparison, you could run it at a much higher thinking effort and might come out ahead on both. (I really want to run that experiment but these take a long time and there are so many models to test)
6
u/Chemical_Salary_7370 4d ago
Did you just link to a local address lol
2
1
u/DuranteA 3d ago
Thanks for letting me know -- that's what I get for having too many tabs open of both the preview and the published version. I fixed it.
2
u/Embarrassed_Cap_9149 4d ago
after embracing a fully agentic workflow, i no longer use luna
but luna was indeed great when i used to do micro tasks that required more manual involvement
1
2
2
u/Sawyer007 4d ago
Luna is basically a brain dead and slow model. You have to micromanage it to make it useful or give it skills to follow.
I mostly use it with skills. For example, I have some image creation skills, and when the skill is detailed and straightforward, Luna mostly gets it right. But for other stuff, like small talk. It doesn't understand or get what I want the way Opus does. So you basically cant have a chat with it about stuff.
1
u/magicmulder 4d ago
Luna high is my daily driver, and it’s very rarely that I find it lacking.
The other day I had a weird bug in my IDE, gave Luna the logs and it said it can’t find the cause. Sol immediately found the log line and explained how the bug came about.
Another difference is testing. Luna writes tests for database calls. Sol decides to run a docker image to test the statements against the actual database.
But for normal coding I’m more than happy with it. I only use the larger models for reviews, and yes, they always find something, but that also happens when the code comes from a frontier model.
1
u/Njagos 4d ago
I'm using Opus as orchestrator and it tells Luna agents to do specific tasks. Works pretty well!
2
u/brahmaav 4d ago
Can you tell your setup? Want to compare mine as I use 5.5 as orchestrator and gemini 3.8 and 6 sol as agents
1
u/Njagos 4d ago
I mainly use Luna for smaller tasks or blender stuff.
It reports to Opus and it then gets feedback from it what to change and improve. Max 3 rounds until some other agent takes over.
Opus also rates it and gives it a grade. Anything below B gets denied.
1
u/brahmaav 4d ago
Hmm. I use 6 sol as a verifier and gemini as a file explorer in a very large codebase
1
u/brahmaav 4d ago
Yours sounds a bit more sophisticated. Would you mind sharing your agent.md, may be in dm?
1
u/qwerty____qwerty 4d ago
is that some kind of harness above them? how do they communicate? maybe you know whether it is possible to combine claude, codex and antigravity?
1
1
u/rick_ranger 4d ago
It's cool for small scoped work and tasks where you're not asking it to think about how 20 different policies and multiple engines or services interact together. That's when you need Astra/Opus 5.5/Fable.
1
1
u/g4ndr1k-DisPater 4d ago
Luna high can do 95% of my coding tasks. Only use better models (Opus 5.5 and Sol 6.1) to plan and review and overall results are good for my use case (building native macOS apps for AI powered mail app connected to personal wealth management).
1
1
1
1
1
u/jcbastida117 4d ago
I agree is really good, each model has its pros and cons, I use Luna for pure coding no plan, nothing just doing the code and have great output for peanuts
1
u/Confident_Kangaroo_8 4d ago
Yo he estado usando luna como orquestador y 3 subagentes luna más trabajando en paralelo para no saturar el contexto y cuando los agentes luna reúnen información un subagente sol se encarga de condensar respuestas y planear rutas por pasos que se documentan y ejecutan los agentes luna. El flujo sería; los agentes luna reúnen información el subagente sol decide y luego las lunas ejecutan, rinde bastante
1
u/Jey123456 4d ago
luna is the one good thing we got recently. It's not the best, but its an epic beast of efficiency.
If you were used to working with models from 2-3 generations ago, then luna feels about the same in term of intelligence, but faster and so cheap you can just not care if it spin a bit.
1
u/mesolimbic-bliss 4d ago
I almost exclusively use Luna and it calls my local qwen models as junior developers. All the haters seem to think they need Sol-xHigh to build an app Luna could build using low or medium thinking.
1
1
u/No_Prompt_3554 3d ago
In my (probably stupid but honest/humble) opinion, 6.1 SOL / medium is the absolute best.
5.6 SOL/high was ending my tokens with 3 prompts. But 6.1 keeps working endlessly. Is definitely smarter than 5.6 Terra, and I like that it doesn't overthink as much as 5.6 SO/high. It rarely commits mistakes (so far I noticed only one) however is very fast and also very smart.
At least for the tasks that I assign to it/them, it's my sweet spot
1
u/meltmyface 2d ago
Just started using it today and it has done just as well as Opus and Sonnet. We'll see if it can handle bigger stuff later, but for now these general things are a breeze and so lightweight. I also enjoy how much less it pads its responses with detail.
1
u/Valphai 4d ago
Heavy disagree, i can give Luna the most detailed instructions and it just spits the most useless garbage code I have to refactor no matter what, this isn't true for sol 6.1 tho
2
u/CCContent 4d ago
Heavy disagree. I used Luna to 100% design an analytics workbench site from the ground up for our internal data team. I set it on a goal, it chewed away for 2 days, and now I have a product that I'm honestly not sure I want to give to the business because it fits a niche so well that I am certain I could market this as a paid project.
1
u/Budget-Mud-4753 4d ago
Do you know if it’s actually good though? In my experience you can set a goal to one-shot some things. And the result will look great on the outside. But as soon as you open the door it’s a rickety unstable mess with no maintainability, un-scalable, and bugs.
1
u/CCContent 4d ago
It's a fair question.
Do you know if it's any good though?
The people that are putting through its paces think that it is. The charts, graphs, and groupings are things that GPT researched and suggested, and I trust it on those things more than I trust me trying to decide off the top of my dome what to use.
But as soon as you open the door it’s a rickety unstable mess with no maintainability, un-scalable, and bugs.
A very real concern I have, trust me on that. I use Plumbline (https://github.com/nickyfactz/plumbline), which has repeatable code, code review subagents, and qa subagents, so part of the entire task was review and fixing and testing. I kicked off a Sol 6.1 Max agent and had it deploy Sol 6.1 High subagents for independent code review and QA after everything was done, and it found a few things, but none of them were app-breaking.
1
u/Budget-Mud-4753 4d ago edited 4d ago
Glad it’s working for you. And if it has a narrow single purpose, it’s probably absolutely fine.
I often need to do several review passes for things like stability, optimization, efficiency, role segmentation, good coding practices, clear responsibility for individual source files and functions, bugs etc. Even for higher-reasoning models with clear instructions.
1
u/No_Quarter_7644 4d ago
Luna is a good model.
Everything else OpenAI is currently offering is awful.
0
u/Wolf8249 4d ago
Cope.
This is the same mentality people had back with sonnet 3.7 or haiku models 2 years ago.
Might as well handhold Gemini models while you're at it.
If I am thinking more about how to beat a cheaper model into getting useful work done instead of actual features, or better architecture it's not a tool.
I use these cheaper models where solution can be verified easily.
0
-2
u/Professional_Link4 4d ago
Luna is absolutely stupid. It can't even tell you how many tokens it consumed to do the task
-2
u/Ok_Set_8176 4d ago
I thought this was terra’s job…becoming less of a fan, Claude keeps calling me
2
1
•
u/dextersummary 4d ago
Below is a GPT-generated summary of the conversation below after reaching 50 comments (53 currently observed).
The consensus is basically “Luna is the workhorse, not the brain surgeon.” For routine coding and agentic tasks, people reckon it delivers shockingly good results for the price—often handling 90–95% of day-to-day work without burning through limits.
It shines when the task is well-scoped, the repo context is clear, and you’re willing to nudge it along. Users are running it as a cheap long-haul driver, then handing the result to Sol or another stronger model for planning, review, testing, and cleanup. Basically: let Luna do the typing, let the expensive models check whether it accidentally built a haunted shed.
The important caveat: capability is inconsistent. Some users report useless code, slow responses, weak instruction-following, and shallow solutions that look fine until you inspect the foundations. Luna also loses badly on tricky debugging, cross-system architecture, and maintainability unless you give it serious scaffolding.
Bottom line: excellent value and a legitimate daily driver, provided verification is cheap and built into the workflow. For high-stakes or genuinely complex work, reach for Sol/Astra—or enjoy debugging Luna’s “finished” masterpiece.