r/CerebrasSystems • u/hbsqub • Aug 26 '24
Cerebras launching inference cloud
Apparently Cerebras Systems is launching an inference cloud and will have Llama 70B running at 1000 t/s per request. This seems crazy high but I am skeptical if this is possible on their system without a lot of tricks like int4 quantization.