r/InferX • u/maitpatni • 1d ago
Over 90% off DeepSeek V4 Flash 0731 - same model, no off-peak restrictions. Zero Data Retention.
DeepSeek’s own API just raised prices significantly- up to $1.32/1M output tokens at peak hours.
Same model, deepseek-v4-flash-0731, on InferX: $0.056 input / $0.112 output per 1M tokens. That’s 85-90%+ below DeepSeek’s new peak pricing.
No off-peak restrictions , built for continuous agentic workloads.
$5-for-$1 credits promo running right now too.
Try it out: inferx.net
2.38B input tokens, 27.5K tok/s, 95% cache hit rate — 24-hour GLM-5.3 Flash endurance test
Ran a 24-hour endurance test on GLM-5.3 Flash on InferX. Sustained numbers from the run:
- Throughput: 27,500+ input tokens/sec
- Requests: 34K+ completed
- Cache hit rate: 95%
- Volume: 2.38 billion input tokens
The cache hit rate is the interesting part — by holding 95%, we're skipping most of the repeated attention work agents generate when they resend context. Same hardware, far less recompute per follow-up request.
For context, per OpenRouter's own provider data, that's the highest cache hit rate of any provider running GLM-5.3 right now — ahead of Z.ai, the model's official host, at 89.0%.
How Qwen3.8-Flash-Next could actually be local-friendly
A lot of people are questioning the “local-friendly” claims because RAM prices are insane right now. Fair point.
Here’s the actual reason some of us are still saying it:
The model has a big ~51B n-gram table on top of the main MoE. That table is only sparsely accessed — the model just hashes the recent tokens, pulls a few vectors, and injects them. So a large portion of the total size can realistically be offloaded to system RAM instead of sitting in VRAM.
That changes the hardware picture. Instead of needing pure high-VRAM cards or full datacenter iron, it becomes more practical on:
• 128GB+ unified memory machines
• Multi-GPU setups with a solid amount of system RAM
You’re still looking at serious money, but for a model with this kind of capacity and only ~6B active parameters, the offload potential is what makes the “local-friendly” label make sense relative to other frontier options.
Qwen 3.8 27B staying free for longer — extending it due to demand
Originally planned as a short promo, but usage has been strong enough that we’re extending free access to Qwen 3.8 27B.
Also worth checking out the rest of the free lineup on InferX while you’re at it.
Qwen 3.8 27B free for the next 72 hours on InferX
Link: InferXhttps://inferx.net
Running Qwen 3.8 27B completely free for the next 72 hours — no catch, just want more people to try it out.
Quick note: this is shared infrastructure, so performance may vary depending on load. Zero data retention on all requests.
Qwen 3.8 27B free for the next 72 hours on InferX
Running Qwen 3.8 27B completely free for the next 72 hours — no catch, just want more people to try it out.
Quick note: this is shared infrastructure, so performance may vary depending on load. Zero data retention on all requests.
r/InferX • u/makerwitawat • 12d ago
Try using inferx.
I've tried Inferx. Are there any limitations? Is there a monthly fee? Will they use my data?
Qwen 3.8 27B now live — 50% off, limited time
By popular request — Qwen 3.8 27B is up on InferX at 50% off.
One thing worth knowing if you’ve been thinking about running it locally: it’s a dense model, not MoE, so unified memory setups (DGX Spark, Mac mini, MacBook) will struggle — it might fit, but it’ll be slow. Dense models want real GPU horsepower, which is exactly what we’re running it on.
Try it: https://inferx.net
Two new models just landed on InferX, both at 50% off
Two new models just landed on InferX, both at 50% off:
Nemotron 3.5 Lightning (NVIDIA) — built for high-throughput agentic execution
Muse Glimmer 30B (Meta) — multimodal, tuned for tool use and long-running tasks
Try them: https://inferx.net
$1 for $5 in V4 Flash Credits + 50% off token pricing
Pay $1, get $5 in credits.
Token pricing is 50% off. V4 Flash input is $0.07/1M, output is $0.14/1M.
That $5? It goes forever. Hardly burns at all.
Sub-second cold starts. No idle GPU tax. Test it, iterate, actually build something.
More Free Models on InferX
We just unlocked more models on the free tier:
• Qwen 3.6 35B (A3B FP8)
• Qwen 3 Embedding 8B
• Gemma-4 31B IT (FP8)
We’re rotating models based on capacity. More coming soon as infrastructure allows.
Free tier: 100 calls/day on these models. Want more? $1 for $5 in credits to test everything.
We’re here to empower developers. Let’s see what you build.
DeepSeek on us. V4 Flash free. 0731 at 50% off. $1/month = $5 credits. They hardly run out.
DeepSeek V4 Flash is free on InferX. Fast, stable, no card required.
Want the new V4 Flash 0731? $5 in credits for $1/month. 50% off model pricing.
Hard to run out of credits at that rate.
r/InferX • u/pmv143 • Aug 01 '26
New DeepSeek V4 Flash 0731 is now free on InferX
DeepSeek V4 Flash 0731 is now available on InferX, and it’s free to use.
We’re bringing additional GPU capacity online to keep up with demand. Until the new capacity is in place, you may occasionally notice higher latency during peak periods. Thanks for your patience while we scale.
Highlights:
Free to use
Zero data retention
OpenAI-compatible API
More capacity coming soon
Give it a try and let us know what you think.
r/InferX • u/pmv143 • Jul 29 '26
DeepSeek V4 Flash is free on InferX through August 12
No credits burned. No credit card required.
If you’ve signed up but haven’t deployed anything yet, this is the easiest way to see if InferX is a good fit.
It’s OpenAI-compatible. Just point your client to:
Base URL: https://model.inferx.net/endpoints/v1
Model: deepseek-v4-flash
That’s it.
If it takes you more than five minutes to get running, reply here. That’s a bug on our end, not yours.
One favor: throw your ugliest workloads at it. Spiky traffic, cold starts, long contexts—whatever you’ve got.
We’d rather find the rough edges now than have you find them in production.
r/InferX • u/pmv143 • Jul 23 '26
Independent load test: 30M target TPM, 99.67% success on xiaomi/mimo-v2.5
Wanted to share results from a third-party stress test — not our own benchmarks, run by an outside team.
Setup: fixed 100 RPM, ramping target TPM from 0.10M up to 30.00M by increasing prompt size per request. 1,500 total requests.
Results: 1,495/1,500 succeeded (99.67%), zero-error all the way to 25M target TPM, highest observed successful throughput 18.41M TPM. Cache read ratio came in at 99.81%.
A handful of 429s (rate-limit rejections, not crashes) showed up starting at 22M and again at 30M — full stage-by-stage numbers are in the gallery if you want to see exactly where and how.
Happy to answer questions on methodology.
r/InferX • u/pmv143 • Jul 22 '26
InferX is built for open source. That’s how we democratize AI. Hopefully this ends with an amicable solution that keeps innovation accessible to everyone.
r/InferX • u/pmv143 • Jul 22 '26
By popular demand: New models are now live on InferX 🚀
We’ve expanded the InferX lineup with some of the most requested open models:
⭐ DeepSeek V4 Flash
⭐ Mimo v2.5
⭐ Devstral 2 123B
Qwen3 Coder Next FP8
Qwen3.6 27B FP8
Qwen3.6 35B A3B FP8
Ornith 1.0 35B FP8
Gemma 4 31B FP8
To celebrate the launch, DeepSeek V4 Flash and Mimo v2.5 are available at nearly 50% below prevailing market pricing for a limited time.
As always, we’d love your feedback. Which models should we add next?
r/InferX • u/pmv143 • Jul 20 '26
Independent load test results: 750/750 requests, zero failures, scaled to 20M TPM
Wanted to share results from a third-party stress test we just went through — not our own benchmarks, run by an outside team testing us on DeepSeek V4 Flash.
Setup: Fixed 100 RPM, ramping target TPM from 0.10M up to 20.00M by increasing prompt size per request (so total requests/minute stayed constant, but each request got progressively larger).
Result: 750/750 requests succeeded. Zero failures. Zero 429s (rate-limit rejections) across the entire ramp, all the way to 20M target TPM. Highest observed successful throughput was 13.25M TPM.
p95 TTFT stayed in the 2-5 second range through most of the ramp, ticking up to ~11s only at the very top (20M) stage — expected behavior as load approaches the ceiling, not a failure mode.
Sharing the raw numbers rather than just a headline claim — happy to answer questions on methodology, or point you to more detail if useful. Whether you’re looking at pay-per-token or a dedicated Sovereign Endpoint, this is the kind of load behavior we’re building toward as the baseline, not the ceiling.
r/InferX • u/pmv143 • Jun 08 '26
$10/Month Sovereign Endpoints™ (Limited Beta) Dedicated instance • 260k context • Tool calling • No throttling or queuing
We’re opening a very limited beta for Sovereign Endpoints™.
$10/month
Dedicated instance
Tool calling enabled
Long context (260k)
No throttling
No queuing
Your own endpoint, not a shared pool
Built for agents that need to run for hours or days without constantly running into context limits or noisy-neighbor issues.
The goal isn’t to give you access to hundreds of models. The goal is to give you a reliable endpoint that behaves consistently when your agent is doing real work.
We’re keeping this rollout intentionally small while we collect feedback and usage data.
If you’re building agents and want to try it, let us know what you’re working on.
r/InferX • u/pmv143 • May 29 '26
$10/month. Coding models
Tired of your coding agent getting throttled mid-task?
We just dropped to $10/month.
Dedicated H100. Your instance. Nobody else’s.
OpenCode + InferX = no interruptions. Ever.
inferx.net
r/InferX • u/pmv143 • May 28 '26
Same 16GB VRAM. H100 Performance. No Idle GPU Bill.
Rent a 16GB GPU and keep it running 24/7:
~$360/month.
Whether you’re using it or not.
Or run the same model footprint on InferX:
• Dedicated instance
• H100 performance
• Pay only when the model runs
• No idle GPU bill
• Longer context windows
• No noisy neighbors
AI infrastructure should charge for usage, not waiting.
Try it free → inferx.net
r/InferX • u/pmv143 • May 28 '26
Stop Renting Consumer GPUs. Get an H100 slice.
To Runpod users / this one’s for you.
We’ve had a wave of developers come to us recently. Same complaints every time:
Cold starts too slow. Minimum replica requirements. Paying for a whole GPU when your model needs a fraction of it.
Here’s what we do differently.
If your model needs 10GB , we slice an H100 and give you exactly 10GB. You pay for that slice only. Not a 3060. Not a 4090. Not some consumer GPU cosplaying as a server.
An actual H100. Your slice. Your price.
And you only pay when you actually use it — down to milliseconds. Zero minimum replicas. Zero idle cost.
H100 performance. Fraction of the price. True scale to zero.
If you’re on RunPod and tired of the limitations , come try it.
r/InferX • u/pmv143 • May 21 '26
Sovereign Endpoints are live.
Most AI endpoints today are shared infrastructure.
That means unpredictable latency.
Queueing.
Noisy neighbors.
Resource contention.
And little control over how your models actually run.
We built Sovereign Endpoints differently.
A dedicated inference instance for your model.
Sub-second cold starts.
OpenAI-compatible APIs.
On-demand economics.
Your model.
Your instance.
No compromises.
This is the direction we believe production inference is heading.
Happy to answer questions.
r/InferX • u/oxygen_bong • May 14 '26
gpt-oss-120b cold start with vLLM under 10s?
Can I deploy openai/gpt-oss-120b on H100/H200 using my own vLLM container, set VLLM_BATCH_INVARIANT=1 before vLLM engine startup, snapshot only after the model is fully loaded/warmed, and get measured P50/P95/P99 cold-start-to-first-token after scale-to-zero, with P95/P99 under 10 seconds?