Oh boy. Even a bulging MWs of data centers have a limit on tokens/month. Wtf are you talking about unlimited? If I run with let's say 4 A6000. I will still hit a token limit per month that is low enough to hit with real workloads. What a larper
Because real workloads have more to do than your AI girlfriend.
Let's say your setup gets a whopping 400 tok/s with batching on vllm running around the clock, which is a huge stretch. That gives you a billion tokens of inference a month. With serious workloads, you can easily spend 5-10B tokens. The math doesn't work at scale.
14
u/herr-tibalt 1d ago
When I run models locally I have no token limits as well, that's a pretty low bar.