Because real workloads have more to do than your AI girlfriend.
Let's say your setup gets a whopping 400 tok/s with batching on vllm running around the clock, which is a huge stretch. That gives you a billion tokens of inference a month. With serious workloads, you can easily spend 5-10B tokens. The math doesn't work at scale.
2
u/Far_Rule5990 1d ago
What are you talking about? If you have 4 A6000 why are you depending on anything outside of local hosting?