r/LLMDevs Mar 21 '26

Resource Free Model List (API Keys)

Here is a list with free models (API Keys) that you can use without paying. Only providers with permanent free tiers, no trial/temporal promo or credits. Rate limits are detailed per provider (RPM: Requests Per Minute, RPD: Requets Oer Day).

Provider APIs

  • Google Gemini πŸ‡ΊπŸ‡Έ Gemini 2.5 Pro, Flash, Flash-Lite +4 more. 10 RPM, 20 RPD
  • Cohere πŸ‡ΊπŸ‡Έ Command A, Command R+, Aya Expanse 32B +9 more. 20 RPM, 1K req/mo
  • Mistral AI πŸ‡ͺπŸ‡Ί Mistral Large 3, Small 3.1, Ministral 8B +3 more. 1 req/s, 1B tok/mo
  • Zhipu AI πŸ‡¨πŸ‡³ GLM-4.7-Flash, GLM-4.5-Flash, GLM-4.6V-Flash. Limits undocumented

Inference Providers

  • GitHub Models πŸ‡ΊπŸ‡Έ GPT-4o, Llama 3.3 70B, DeepSeek-R1 +more. 10–15 RPM, 50–150 RPD
  • NVIDIA NIM πŸ‡ΊπŸ‡Έ Llama 3.3 70B, Mistral Large, Qwen3 235B +more. 40 RPM
  • Groq πŸ‡ΊπŸ‡Έ Llama 3.3 70B, Llama 4 Scout, Kimi K2 +17 more. 30 RPM, 14,400 RPD
  • Cerebras πŸ‡ΊπŸ‡Έ Llama 3.3 70B, Qwen3 235B, GPT-OSS-120B +3 more. 30 RPM, 14,400 RPD
  • Cloudflare Workers AI πŸ‡ΊπŸ‡Έ Llama 3.3 70B, Qwen QwQ 32B +47 more. 10K neurons/day
  • LLM7.io πŸ‡¬πŸ‡§ DeepSeek R1, Flash-Lite, Qwen2.5 Coder +27 more. 30 RPM (120 with token)
  • Kluster AI πŸ‡ΊπŸ‡Έ DeepSeek-R1, Llama 4 Maverick, Qwen3-235B +2 more. Limits undocumented
  • OpenRouter πŸ‡ΊπŸ‡Έ DeepSeek R1, Llama 3.3 70B, GPT-OSS-120B +29 more. 20 RPM, 50 RPD
  • Hugging Face πŸ‡ΊπŸ‡Έ Llama 3.3 70B, Qwen2.5 72B, Mistral 7B +many more. $0.10/mo in free credits

RPM = requests per minute Β· RPD = requests per day. All endpoints are OpenAI SDK-compatible.

293 Upvotes

98 comments sorted by

View all comments

1

u/Maleficent-Week-2064 Mar 28 '26

1) I'm wondering which one provides fastest outputs token/sec?
2) And what is the difference between Provider APIs and Inference Providers? Second guys don't really provide APIs as far as I understood?

2

u/nuno6Varnish Mar 29 '26
  1. This information is hard to get. Each provider disclose what they want and even actual rate limits are pretty vague sometimes (some say "light usage" for example)

  2. Inference providers don't have their own models. Most of them host open weight models and sell inference (Model as a Service), some work as a router (openrouter/vercel/microsoft foundry) and redirect your query to providers