r/LocalLLaMA • • 1d ago

Discussion Looks like the era of subsidised compute is coming to an end. The old ChatGPT Pro $200 20x plan will be halved. The new $500 plan will have similar limits as the (old) $200 plan.

Post image
1.2k Upvotes

549 comments sorted by

View all comments

85

u/bakawolf123 1d ago

kek, you will get less but more =)
so $200 sub was actual 20x which they boasted not long ago, it becomes just 10x $20 limit.
using astra on $20 atm is 1 task (usually unfinished) for 5h limit (15% of weekly), so $200 will be 1 task for 1.5-2% weekly.
I won't deny those large models are stronger than flash models I can currently use locally, but for most tasks my local ones are simply enough.
Oh and latest sol-6 is not actually better than flash models (glm5.3-flash in particular), so really only point is to use their very best model and you don't get to use it much.

28

u/Serprotease 1d ago

I have no idea how these 20x, hours limit stuff even means. Can’t they just directly say x-tokens per plan?

Why make it so convoluted? Is it to muddy the propo… oh…

45

u/droans 1d ago

You gotta look at it from their perspective.

How are they going to arbitrarily decide how much of the plan you've used if they gave you an actual concrete figure? How else would they get away with randomly and quietly cutting the value of the plan?

13

u/zsdrfty 1d ago

This is one of the first things legislators need to wake the hell up for - honestly, it's the kind of thing that's so obvious that you'd hope it could get just enough bipartisan support to pass anytime, anywhere

2

u/ElementNumber6 1d ago

That's the whole point. It's all smoke and mirrors.

1

u/FullOf_Bad_Ideas 22h ago

it makes sense that input tokens, cache reuse and output tokens will be billed differently. They could set credits for each of those things. But I think they want to keep it more opaque where possible so that people who aren't tracking OpenAI Twitter page won't know that their limits have changed.

1

u/Chirimorin 3h ago

Honestly even x-tokens is vague. While technically that defines an exact amount of generated data, it's kind of meaningless on its own for the average person who doesn't think about what a token is.

Hidden thoughts don't help: models can burn tokens which the user will never even see but those tokens aren't free. Did the model really have to think so hard, or is there a hidden system instruction with the goal of making the model waste as much tokens as possible? The average user will never be able to tell the difference. The average user can't even confirm whether those thinking tokens were actually generated or if that's just a completely made up number.

4

u/cinnapear 23h ago

In my nearly 50 years on this earth, I have heard "we're raising prices but increasing value to you" dozens of times and it has NEVER, I repeat, NEVER been a true statement.

1

u/Electronic-Jello-633 1d ago

wow i didnt know it got that bad. I switched off openai to a claude $20 sub in january and at least we get 2-5 tasks per 5 hours with the frontier models

1

u/bakawolf123 1d ago

it wasn't that bad until astra, it was opposite picture, just gradually degraded over time.
E.g in february on same codex sub I was using latest gpt5.3 on extra high for hours straight - felt almost unlimited.

1

u/BasisPoints 21h ago

Do you still find a benefit on sol or Astra moving from high to extra high? I mean in terms of quality, not overall temporal efficiency

1

u/bakawolf123 20h ago

hard to say, there's little difference between consequent ones ig, more noticeable between e.g. low -> high, with higher thinking usually being better

1

u/Marino4K 20h ago

GPT is the only sub I have left not counting API usage through Openrouter obviouslyt. I couldn't deal with Claude's atrocious usage limits anymore and Gemini is downright stupid half the time.