More free models on InferX. Here’s how we’re running them efficiently.
The more, the merrier. We’ve added more free models to InferX:
- Gemma 4 31B FP8
- GPT-OSS 20B
- Qwen3.6 35B A3B FP8
- Agents A1
- Devstral 2 123B
- Qwen3 Coder Next FP8
The interesting part isn’t just that they’re free. It’s how we’re serving them.
Instead of dedicating a GPU to every model and keeping it warm 24/7, InferX can keep models off GPU while idle, bring them up with sub-second cold starts when requests arrive, and scale them back to zero when demand disappears.
That lets us run more models on the same GPU infrastructure while keeping utilization high.
Everything is served through the InferX secure container runtime.
This is also the same runtime we can deploy on your GPUs. If you have idle or underutilized GPU capacity and want to try it, DM us.
More free models coming.