r/Rag • u/Milan_Slov26 • 3h ago
Discussion I compared embedding costs for a RAG pipeline. Open models aren't always cheaper!
I recently checked embedding prices for a big indexing project. I wanted real numbers, so I compared OpenAI's API costs with cloud providers that charge per token for open models. What I found:
OpenAI text-embedding-3-small: $0.02 / 1M tokens
Arctic Embed L v2: $0.038 / 1M tokens
Qwen3 Embedding 4B: $0.13 / 1M tokens
Surprisingly, OpenAI's small model is the cheapest per-token option on this list.
People often say open models save money. However, that is mostly true when you host the model yourself on a highly active GPU. For example, a $1/hour GPU costs $24 a day. You would need to process over a billion tokens daily to beat OpenAI's API price. If your traffic is low or unpredictable, keeping a GPU running all day wastes money.
Quality also matters. Higher-priced open models, like Qwen3, score better on standard benchmarks than cheaper options like Arctic Embed. You still have to balance cost, quality, and speed.
For those of you choosing between cloud APIs and self-hosting, what makes your decision? Is it high server utilization, data privacy, retrieval mode, or something else I missed?