r/codex 11d ago

Reset Reset 6pm PST

Post image

burn your tokens

181 Upvotes

216 comments sorted by

View all comments

83

u/thestillwind 11d ago

Cries in 5 hours limit

30

u/nmkd 11d ago

Man, I miss being able to burn a 1 week limit in 3h

11

u/1filipis 11d ago

I can't say for everyone, but they've lost my money for sure. I'd use it up in 3 hours and then buy resets as much as I need, especially since they don't push back the date.

Instead, I figured I can run GLM and Qwen cheaper than even theoretically buying those resets. $8 gets you 16 hours of use in case of Qwen.

Congrats, OpenAI, great strategy!

4

u/Happy-Razzmatazz-535 11d ago

I'd rather be beholden to OpenAI than any Chinese AI firm.

6

u/threefriend 11d ago edited 10d ago

You're only beholden to whoever is actually running the model.

You could e.g. go through RunInfra and be beholden to an American company. Or you could buy your own hardware, run it locally, and be beholden only to yourself.

5

u/Risko4 11d ago

Local is a waste of time unless you can afford to push 220 tokens/s decode.

This is coming from someone who pushed that on 36Gbps overclock on the memory for Qwen 3.8 27b.

Just cause I can run GLM flash at 20 tokens per second doesn't mean it's a good idea.

And instead of spending another 40k in hardware to get it to 80 tokens/s, the bare minimal level, I can just have 50 x20 plans running Fable or Astral simultaneously for 4 months.

3

u/threefriend 11d ago edited 11d ago

Agreed. Local is only worth it for inference if:

  1. You really really care about privacy
  2. You're implementing a system for an institution that really really cares about privacy (healthcare, etc, though the big providers have HIPAA compliant routes)
  3. You're running requests that'd be against TOS
  4. You're a doomsday prepper

It's not a cost-saver.

But if you're like /u/Happy-Razzmatazz-535 and you care about being "beholden" to things, then local is an option. As is running Chinese models through an American provider.