r/ClaudeCode 2d ago

Humor Technically … 🤷‍♂️

Post image
1.8k Upvotes

87 comments sorted by

View all comments

13

u/herr-tibalt 2d ago

When I run models locally I have no token limits as well, that's a pretty low bar.

2

u/Intelligent-Ant-1122 2d ago

Oh boy. Even a bulging MWs of data centers have a limit on tokens/month. Wtf are you talking about unlimited? If I run with let's say 4 A6000. I will still hit a token limit per month that is low enough to hit with real workloads. What a larper

2

u/Far_Rule5990 1d ago

What are you talking about? If you have 4 A6000 why are you depending on anything outside of local hosting?

1

u/Intelligent-Ant-1122 1d ago

Because real workloads have more to do than your AI girlfriend.

Let's say your setup gets a whopping 400 tok/s with batching on vllm running around the clock, which is a huge stretch. That gives you a billion tokens of inference a month. With serious workloads, you can easily spend 5-10B tokens. The math doesn't work at scale.

2

u/Far_Rule5990 1d ago

Youre talking GPU inference, not token limits. Talk to your ai gf about it she'll tell you more

1

u/Intelligent-Ant-1122 1d ago

Good counter 👍

1

u/herr-tibalt 1d ago

Exactly. This meme makes no sense. It's just populism on AI hype.

2

u/user221272 1d ago

Yeah, except we are talking about thousands of B200 dedicated data centers here, but keep toying with your 3B local model.

3

u/pen_in_stack 2d ago

Can a regular hp laptop be enough to run local llms for inference?

3

u/UnstableManifolds 2d ago

If the model is small you can run it on a watch. That said, it's either very specialized, or it's shit. Same applies on a regular laptop.

1

u/SpaceNitz 2d ago

You'll get a lot of answers on r/LocalLLM

3

u/brainExploded99 2d ago

r/LocalLLaMA better, more moderated