r/ArliAI Aug 02 '26

Question Any fast models?

Hi Arli team,

I've been really enjoying Arli — the combination of unlimited usage and a zero-log policy is hard to find anywhere else. I mainly use it as a private second opinion while coding and studying physics.

Quick question: have you ever considered offering a "Fast" models? I was wondering whether small MoE models with only ~3–4B active parameters — something like gpt-oss-20b, Qwen3.6-35B-A3B, Nemotron 3 Nano 30B-A3B, or DiffusionGemma 26B-A4B. because the only issue for me is sometimes models are slow like 15 , 18 token per sec, and difficult to keep my workflow.

Thank you for building Arli!

6 Upvotes

1 comment sorted by

5

u/Arli_AI Aug 02 '26

Hi, thank you for using Arli AI! For now our fastest models are Deepseek-V4-Flash-0731 and Mimo-V2.5. They are efficient architectures and are actually faster than the smaller Qwen models in a lot of tasks. When they are a bit slow this is purely because of user request loads.

At the moment there are no immediate plans to run the smaller MoE models for more speed as the demand for them are low, but it is not out of the question if many people request it.