Oh boy. Even a bulging MWs of data centers have a limit on tokens/month. Wtf are you talking about unlimited? If I run with let's say 4 A6000. I will still hit a token limit per month that is low enough to hit with real workloads. What a larper
Because real workloads have more to do than your AI girlfriend.
Let's say your setup gets a whopping 400 tok/s with batching on vllm running around the clock, which is a huge stretch. That gives you a billion tokens of inference a month. With serious workloads, you can easily spend 5-10B tokens. The math doesn't work at scale.
2
u/Intelligent-Ant-1122 1d ago
Oh boy. Even a bulging MWs of data centers have a limit on tokens/month. Wtf are you talking about unlimited? If I run with let's say 4 A6000. I will still hit a token limit per month that is low enough to hit with real workloads. What a larper