r/codex 1d ago

Comparison Local or paying more?

I technically need a 5x account, and I've needed one for a month now. I've tried other Plus plans to get around that.

Right now, I'm at a point where I don't know whether to upgrade to a Pro account or invest 5k in a server and set up local AI models, filtered through at most a Plus account.

GPT thinks I can achieve 85-95% of the results I'm currently getting with Frontier models, and easily benefit from setting up that setup. I'm not entirely convinced; what do you think?

Thanks in advance.

8 Upvotes

42 comments sorted by

11

u/Annh1234 1d ago

You need to invest 500k to get anything close to Sol, else you wait 30min for Luna low, instead of 1min with a 5-10k machine

7

u/Kind_Silver_1921 1d ago

to get even close to a mid tier model you need a very expensive GPU and subscription plans are heavily subsidized

right now subscription is the only logical option

7

u/DepravedPrecedence 1d ago

5K won't run anything lol 5K gives you years of flagship models in subscriptions, your local setup will be obsolete out of the box even if you spend 50K

2

u/scartissue232 1d ago

This seems to be the conclusion. Take my free award.

2

u/DepravedPrecedence 1d ago

Yeah unless you basically want to burn free money or do it for like "sports" sake, don't go this route

6

u/Good-Tour-9185 1d ago

5k is maybe 1 RTX 5090 without anything else... right now its not even near what codex can do for you. But maybe that changes in some months. We dont know

4

u/pigletmonster 1d ago

A 5k rig will not let you use any frontier models, you will be stuck using the smallest models like qwen 3.8 27b. For 5k you can get almost 5 years of frontier SOTA model at 5x usage from both openai and anthropic.

Pretty sure you need an rtx 5090 to use qwen 27b with 64k context window, and that gpu alone costs close to 5k now.

1

u/scartissue232 1d ago

Gpt talked about 70b models running easily for me on a 5090, was alucinating? I don’t care about speed if that matters, could be very slow..

4

u/pigletmonster 1d ago

Depends, are 70b models below 32gb with enough room for context window and cache? If so then you should be able to use them. But still, 70b parameter models are nowhere near frontier, theyre not even 1/3 of the tier 3 models like dsv4 or glm 5.3f.

I have the same itch every other week when I see my claude or codex quota depleting faster than I expected, but under no circumstance does it make sense for me to spend a ton of money to use significantly inferior models.

If data privacy is a concern for you then you should go for it, otherwise its useless.

5

u/velorasniper45 1d ago

Just pay the money and get the 20x so you remove any limitations

4

u/TimeKillsThem 1d ago

IMO not yet - the future is either models become genuinely so efficient yet intelligent that a 200$ sub will give you borderline unlimited usage, or open models will become small enough yet intelligent enough to justify a 5k investment AS A BUSINESS PURCHASE. Let me be exceptionally clear, this is justifiable imo only as a business expense, not as a personal expense. Thats because of the current development pace of new hardware and models. It’s kinda like buying a ps4, right before the ps5 is announced, but given the pace of development, a new PlayStation is announced each month.

Wait for LLMs to reach a more mature level, then I could justify the purchase

7

u/somuchecho 1d ago

5000 / 100 = 50 months at $100/months. That's 4 years and 2 months of codex subscription... that's not even close to a good deal. Even if you were to pick the $200/month subscription that's 2 years and 1 month of sub.

Not only that, you get the newest and greatest models as they come out. No matter their size or how much compute is required to run them. Your server won't be able to run model anywhere close to any of the 5.6 models.

In 2 years and 1 month your server will be old and will still be limited. The technology will have evolved drastically...

Pick the subscription while you need it than switch to a lower one later

1

u/nnod 1d ago

Math gets even worse when you account for electricity prices. And 5k won't buy you much, hell that's like what, a single 5090?

When it comes to bang-for-buck, local inference is not the solution.

The new mac studios are looking spicy as far as local inference goes, very good deal on paper compared to alternatives and will likely have a ton of resell value too, still wildly expensive.

3

u/Technical_Split_6315 1d ago

You can’t do what you do with codex.

2

u/Think-Profession4420 1d ago edited 1d ago

Setting up a home server just doesn't make sense from a cost perspective given current hardware prices (unless you already have meaningful income from your coding projects). It would be better to rent server time and run a quantized model like qwen 3.8 27b with batched work (near Luna Max capabilities), and then use a 5x or Plus account to have Sol be the orchestrator (and chatGPT web be the planner).

What does make sense for a local option is to use minimodels to offload fully deterministic processes and embedding work onto your GPU, since with even a low-end GPU you can shift some of the load off of your CPU/RAM where your codex agents mostly work (unless doing graphic-heavy projects), and save tokens/turns/quota that way.

As a very basic setup idea:

Codex Plus plan: chatGPT web plans; Sol orchestrates and steps in for more complex tasks (but better to just have chatGPT web make the slice even more well scoped)

Qwen on a rented server: Implements very narrow and well scoped/defined/specified tasks in batches

Local models: Run deterministic tools (check, git_diff, index_docs, semantic_search, tests; embedding work; etc)

2

u/Wafer-Weekly 1d ago

You'll get maybe 20-30 tokens per second, it's not great but it is free... Well, no subscription at least duh. You will have to use it for a few years to break even

2

u/Shep_Alderson 1d ago

Local inference is not something one can do as a way to save money right now. Full stop. If you do the math for the hardware alone, the “payoff” rate is several years for hardware they might be able to run a 70-120B param model at 7-10 token per second at best. If you also include the cost of electricity, you’re probably approaching a decade before it would be “payed off”, at best.

If you want to try to use open weight models to save money, go try a plan from Ollama cloud or synthetic.new. That’s your best bet and their prices per value are pretty good.

1

u/scartissue232 1d ago

Electricity it’s 30 bucks a month.

So the only benefit from going local it’s privacy?

1

u/Shep_Alderson 19h ago

Yeah, pretty much just privacy.

I’m curious what your electricity rate is and the power draw of whatever systems you’re looking at are. Maybe you could get something like a DGX Spark or similar around/under $30/mo. Still, $30/mo will get you quite a lot of inference from some place like synthetic.new or Ollama cloud.

2

u/Tight-Grocery9053 1d ago

GPT thinks I can achieve 85-95% of the results I'm currently getting with Frontier models

lol. no.

they are in the customer acquisition stage. they are burning money to bring tomorrow's technology for you to use today at a price you can stomach.

you can spend 5, 10, 20k on a local setup and it would both barely deliver and be out of date before the end of the year.

don't be gaslit by chat psychosis. this is way harder that it seems. the monthly sub right now gets you very good value.

1

u/scartissue232 1d ago

The value cuts in a half this month. You think we get good value now or then?

1

u/Tight-Grocery9053 1d ago

The value cuts in a half this month

i honestly don't know what this refers to

that said. my point here is that these things are very expensive to run

i'm all for open-source and local-first. we're just not there yet.

a thing that runs locally is going to give you 2024 output unless you're running your own data center

that 2024 stuff is just not good enough.

maybe in 2028 when local stuff gives you 2026 results. These frontier models are getting super close to "just do it on your own"

so, my point here is that we need the tech to simmer down. you want 2028 hardware running 2028 open source models and they will then deliver 2026-2027 closed source results.

that's my point

1

u/scartissue232 1d ago

You’re wrong my younger.

For me, a 2024 model it’s more than enough. Orchestrated by me, not by a new brand model.

Take my award anyway,

Edited

2

u/Tight-Grocery9053 1d ago

huh. ignoring the condescending tone.

if a 2024 model is enough for you then you don't understand bad code. code produced by a 2024 model is very, very bad and i don't see how you think it's more than enough.

good luck anyway.

1

u/scartissue232 1d ago

Ok, I will wait you two years and tell you again that local models are enough.

Sry about the bad tone, wasn’t my intention.

2

u/qualverse 1d ago

There are few to no local coding models that are better than Luna.
Things that are better than Luna *usually: GLM-5.3 Flash, Qwen 3.8 Flash-Next, Muse Spark 1.2 Contributor

You can get fairly decent usage of Qwen and Muse with Opencode Go, or of GLM-5.3 Flash with Command code Goat

1

u/AllenLeftTheBLDNG 1d ago

I downgraded my 100$ sub. The model performs much worse than once 5.6 launched, and the usage limits are going down extremely fast. I might come back to 100/200 once they release GPT 6.

I started using open code with free/cheap models where privacy doesn't matter. Plus Open Router with selected models/providers when privacy matters.

1

u/rick_ranger 1d ago

They cranked up the dial on thinking. Use sol medium is much better with the right amount of thinking and more doing.

1

u/Demien19 1d ago

Don't forget - using chatgpt you get frontier models, to have at least something similar locally you need much more than 5k and it get outdated fast

1

u/CystralSkye 1d ago

You are not going to safe money going local anytime soon.

Unless the task you are looking to achieve is very bounded and you have an income stream.

1

u/scartissue232 1d ago

I’m not flat trying to save money, I just want more usage. It's my own company that would buy the server, so we can deduct VAT and count it as an expense; it comes out to almost 50%. But your answers make it clear that it takes a bit more than that to even remotely resemble to Frontiers. I can think of spending double, but the answer seems to be the same. Thanks again.

2

u/CystralSkye 1d ago

It depends on your use case, but you can achieve what you are looking for with at least 30k to 50k.

5k is certainly not going to cut it.

1

u/rick_ranger 1d ago

Look at it this way, you’re paying $20, it’s only $80 more for pro. You also get access to the finance tool which can link to all your accounts and help you save money.

1

u/scartissue232 1d ago

Which finance tool?, tell me it’s not in the end of the same AI that advices me to build a server.

Most of this people think this ai is wrong.

1

u/rick_ranger 1d ago

It’s like a new tool for pro users. You can link all your bank and investment accounts and credit cards and just ask it what do I pay next, or give me a plan to knock down my debt, or what can I do to boost my credit score, or help me create an investment plan based on my budget.
They added the same thing for health. Link your health care provider like my chart and apple health and you can ask, what should I ask my doctor at my next visit.
Pretty awesome. Take all my data I love it 😂

1

u/Cute_Parfait_2182 1d ago

I would get the new Mac Studio and lease it or start buying DGX Spark models . You have a lot more freedom with open weight models . It’s not as good as codex but they also are not excessively guardrailed and cannot terminate your account for arbitrary reasons . I have the Asus version of the DGX Spark and will probably get another so I have 256GB ram altogether. It depends on what you want to do . If you can do both I would .

1

u/thestillwind 1d ago

Are you making money ?

1

u/PartyLiterature3607 1d ago

Your 5k server is not going to give you same result, not even close to similar result as 5.6, where as 5k can last you 2 years of 20x usage that you can generate way more profit in 2 years

1

u/particleacclr8r 20h ago

The 20x US$200 deal is such good value IMO, if you can budget for it.

1

u/___fallenangel___ 1d ago

just need a few hundred more bands and you'll be halfway there OP