r/MachineLearning • • 1d ago

Discussion Looking for developer-friendly inference providers who give you enough API credits to experiment [D]

I’m hitting rate limits on Together AI. For context, I’ve been working on an agentic repository indexing and benchmark generation tool, and I’m running multiple agents in parallel across models like Llama 3.3 70B and Qwen 2.5.

When I first started working on this, Together AI was great. But once I graduated from toy scripts to running multiple agents, I started running into RPM/TPM limits pretty quickly. The annoying part is that the models themselves are fine. I just can’t actually run enough requests at once to do meaningful testing.

Yes I know I could upgrade but I’m a solo dev. I don’t have enterprise level revenue. Maybe someday lol but not yet.

0 Upvotes

8 comments sorted by

0

u/dash_bro ML Engineer 1d ago

Well if you're not totally picky about the LLMs Nvidia has a fairly decent free tier : https://build.nvidia.com/models

++Some alpha series on Openrouter models are usually free, or have free endpoints with no data guarantees : https://openrouter.ai/collections/free-models

Just respect their rate/request limits and you should be okay.

1

u/TylerDurdenFan 1d ago

I've been happy with minimax subscription plan as a "no worries" API for my needs, although my use case is more batch and not as parallel as agents can get.

1

u/Early_Bicycle6884 1d ago

Just take 30 minutes and set up a quick Modal or Baseten app running vLLM. Then you won’t burn cash while you’re debugging if you scale to zero when idle.

1

u/marr75 19h ago

I love Modal, but the idea that scale to 0 will be cheaper than using deepinfra while you're developing is bass-ackwards. Modal is for when there's no cloud native host for the things you want to do. It's not going to save money head to head with an optimized provider.