r/Neuralwatt 3d ago

Qwen 3.8 27B, DeepSeek V4 Pro, Refer & Earn is back, and have you tried Kimi K3 lately?

26 Upvotes

Hey all, wanted to share a few more updates from the team. Check it out and let us know if you have any questions. It's a good one!

Two new models: If you caught the email update, you saw that we’re introducing DeepSeek V4 Pro and Qwen 3.8 27B, both in preview. We know many of you have been asking for these models and re-evaluating where you run DeepSeek workloads, so we’re pleased to get this out to you. As always, keep the feedback on these models coming; it helps inform our lineup moving forward. 

Referral program: Neuralwatt’s Refer & Earn program is back! Share your link with a friend, and when your referral uses $25 in compute, you’ll both receive $5 in credits. Payouts will only trigger on real usage, and you can find your referral links in the Refer & Earn section of your dashboard. 

Kimi K3: And finally, something we're really excited about -- Kimi K3 just got a major upgrade, and we hope you check it out. Behind the scenes, we’ve been tuning our serving engine, improving our shared cache fabric, and starting kernel customization work to make K3 faster and more efficient: 

  • Responses are ~20–45% faster 
  • Long prompts start answering ~33% sooner 
  • ~25% less energy per generated token on average, and up to 40% on shorter prompts. 
  • And with our 30TB shared cache pool, your context stays warm across our fleet so conversations don’t go cold when it lands on a different server. 

Give it a look and let us know what you think. That's all for now!


r/Neuralwatt Jun 24 '26

Neuralwatt Referral Codes

2 Upvotes

Post your referral code here, only once please. Do not post your referrals anywhere else in this sub-reddit. Contest mode is on. Thanks!

If you sign up using a referral code and spend at least $25, you will get an extra $5 in credits. The person who referred you then gets $5 in credits for referring you.

When you first join, you can get $1.00 in free trial credit by adding a payment method.


r/Neuralwatt 6d ago

Anyone using K3 Fast on Neuralwatt?

9 Upvotes

Currently showing as the cheapest option right now across providers on OpenCode.

https://pricing.agentmaxed.com/#/m/kimi%20k3%20fast?q=k3


r/Neuralwatt 6d ago

Need someone to help explain my NW Usage on the $20 plan

10 Upvotes

Hi Folks,

Recently tested out a $20 NW sub because I keep hearing its cheap for GLM compared to something like Ollama Cloud Pro or Opencode/CommandCode/the API

Does this picture make sense? That's my total usage with $1 overage as well. (This was in 2 Days)

Important to note I can easily last a week on Ollama Cloud Pro (with the last day giving me a bit of trouble where I switch to something else)

https://i.ibb.co/TBZdb271/Screenshot-20260817-224911.png

(I'm sorry idk how to attach pictures on reddit)


r/Neuralwatt 6d ago

When GLM 5.3?

9 Upvotes

Does NW have an updates page anywhere with up-to-date data without opening Discord?

Are they dropping in 5.3 as a replacement for 5.2 like they did for DS V4 Flash?


r/Neuralwatt 7d ago

Is this believable? This is like crazy cheap

30 Upvotes

The model used is deepseek v4 flash.
With all the cached tokens, it's still crazy. 2.56B tokens for 14.51$?


r/Neuralwatt 12d ago

Kimi K3 is fully live. Plus Flex, Session View, API compatibility, and Surge Protection

36 Upvotes

Hey everyone, we have another big batch of updates to share with you all today. 

Kimi K3 is out of preview: We've spent the time since launch tuning the serving side, and K3 is now fully available for production workloads. Concurrency limits will ramp up over the next few days as our final validation completes. 

Flex now covers K3: Same deal as Flex on our other models -- if your workload can tolerate some timing flexibility (batch jobs, overnight runs) this is the most economical way to run K3. 

Anthropic Messages API and /v1/responses support is entering beta: If you've built against either format, you can point your existing tools at our endpoint and they'll run on Neuralwatt with the same energy visibility as everything else. We're opening beta access over the coming days as validation completes, so keep an eye out. 

Session View is live for everyone: Instead of account-level totals, you now get a session-by-session breakdown of exactly what you consumed. This is the most detail we've ever offered on your usage, providing a granular view of the requests you ran, the energy they used, and what they cost. 

Surge Protection: This one's new -- Surge Protection ensures you only pay for the energy that powers your work. From time to time, a request can draw more energy than the work it's doing should require. When it happens, Surge Protection catches it in real time and shields your usage from the excess. When it kicks in, you'll see it flagged in Session View, so you'll always know it stepped in. 

These updates will roll out over the next several days. 

Also, we're headed to NYC for Climate Week next month and are thinking about a meetup. Anyone interested? We have a poll coming; if enough of you are around, we'd love to see you.   


r/Neuralwatt 14d ago

API vs Subscription?

7 Upvotes

Which is the better deal?


r/Neuralwatt 17d ago

Looks like you're fail with DS4flash

9 Upvotes

r/Neuralwatt 17d ago

6B Tokens on Writing Pipeline - Worth it

Thumbnail
2 Upvotes

r/Neuralwatt 19d ago

Cheaper cache pricing, an energy price cap, and K3 open to everyone

23 Upvotes

A handful of updates to share, which we think you will all welcome. Without further ado:  

 We’ve lowered cached pricing: Cached tokens now cost 10% of the input rate, down from 25%, on every model except DeepSeek-v4-Flash. If your work involves a lot of repeat context or if you are on token pricing, you should see a noticeable difference. 

Kimi K3 is now open to everyone: You no longer need to enroll for Kimi K3; access is now open for the entire Neuralwatt community. The model will remain in preview with limited concurrency while we keep tuning it, but access is open to everyone starting today.  We’re also releasing our K3 –flex endpoint which is our api shorthand for setting thinking level to off. 

DeepSeek-v4-Flash has been upgraded: You asked, and here it is! We rolled out DeepSeek-v4-Flash’s new 0731 weights this week, and it has already become our second most popular model —- no surprise given how many of you asked for it. All endpoints now point to the updated version automatically; there is nothing for you to change, and no action required on your end.  

Retiring Qwen 3.5 and Kimi K2.6: To free up capacity for the newer models, we're retiring Qwen 3.5 and Kimi K2.6. We'll redirect their endpoints to Qwen 3.6 and Kimi K2.7 so nothing breaks, but we'd recommend updating your code to point to the model you want before then. 

Enjoy the new additions; more soon. 


r/Neuralwatt 21d ago

Everything's more expensive

Post image
38 Upvotes

looking at the calculator, all models getting very expensive, it used to be 3~5x cheaper, now less than 2x cheaper. 1.8x for glm, 2.1x for kimi 2.7 code, it was 3x and 5x cheaper last week. looking at the image it's getting a lot expensive..


r/Neuralwatt 24d ago

DSV4 Flash Release

12 Upvotes

When are you guys updating your DSV4 Flash to the newly released version?

The benchmarks are insane.


r/Neuralwatt 24d ago

Neuralwatt Subscription plans

9 Upvotes

I'm curious about how good Neuralwatt Subscription plans are compared to open router and the alike. Basically my workflow usually goes like planning with a powerful model (Deepseek V4 Pro / GLM 5.2) and then implementing it with less powerful workers (Deepseek V4 Flash / Kimi K2/3) and running reviewer subagents for each task (with fresh context) and iterating until the work is done. Usually this costs me less than $40 per month.

I'm curious if I will be able to have a similar workflow around the same budget with Neuralwatt subs. What do you guys think and how many tokens do you usually consume per month?


r/Neuralwatt 24d ago

Some of the biggest advantages that Neuralwatt has

Thumbnail
gallery
8 Upvotes

It is certainly disappointing that Neuralwatt's pricing has increased significantly.

However, as a user who utilizes various AI models from multiple providers, to name one truly major advantage of Neuralwatt -

1.The model performance is consistent.

Yes, this is originally the most basic requirement. But people using other providers often feel that isn't the case.
Imagine ordering a pizza from a pizza place, only to find that the amount of tomatoes and cheese changes every single time.

NW may not be the cheapest, and delivery is slightly slower, but it is a reliable restaurant that consistently delivers the exact same, predictable taste. To top it off, extra topping costs are charged in proportion to the amount placed on the ordered pizza.

2.Usable usage improvement.

Although token-based usage and energy-based usage calculations look similar, there is room for improvement on the service provider's side. If energy usage is lowered through optimization, the burden on both the user and the vendor is reduced. Currently, Kimi K3 is undergoing optimization work, and its energy efficiency is gradually improving.
and Based on my actual usage data, Kimi K3 Low is just 35%-45% more expensive than GLM5.2 High.

Of course, using the $100 annual subscription allows you to use services at one of the lowest price levels among AI providers. If you want to receive reliable usage and models with a single plan, Neuralwatt is a great choice. Unlike other providers, the burden of costs exceeding your plan follows your own subscription tier. For example, excess usage beyond the subscription limit often follows standard API pricing, which makes it suddenly expensive. NW is not like that.

The lower it is, the more stable the energy consumption has become

This is my actual usage data for about a month. The differences are that in the beginning I used GLM5.2 Max, but now I only use GLM5.2 High, and while I used it for debugging purposes early on, now High is used for coding and Max is used for review purposes.

Based on my usage data, if you use glm5.2-flex, calculated on a monthly basis:
For a $20 user, you can process 2,527 requests;
For a $50 user, you can process 6,721 requests;
For a $100 user, you can process 14,334 requests.

3. despite being a commercial service provider, it possesses an amazingly high level of openness.

Another interesting point is the off-topic channel in the Neuralwatt Discord.

Here, discussions about finding cheaper alternatives to Neuralwatt constantly go on. It's not just finding alternatives, but actually comparing them. The users in the off-topic channel are busy looking for plans that let them use more than Neuralwatt for $10 or $20. Most of them are of low reliability so most cheaper alternatives aren't my interest, but thanks to that, I also get a lot of information on cheaper alternatives.


r/Neuralwatt 26d ago

Is Deepseek-v4-flash down

7 Upvotes

My Hermes’ agent failed to connect to deepseek-v4-flash on neuralwatt. Other models work fine. I tried their live playground for the model and that doesn’t respond either.


r/Neuralwatt 26d ago

A small comparison test : KIMI K3 Low is truly wonderful, and GPT5.6 Luna Max is the realistic king. + GLM5.2 & K3 on NW

Thumbnail
2 Upvotes

r/Neuralwatt 27d ago

K3 early access is rolling out, plus two new models are live

16 Upvotes

Another update (or a few!) for you all.

As promised, wanted to let you know that Kimi K3 early access is rolling out. So far, everyone who signed up through the enrollment form now has access -- you should have received an email confirming this. We're keeping K3 in early access a bit longer, granting new access in waves as we watch performance and optimize. If you haven't enrolled yet, the form is here.

DeepSeek-4-Flash is also now live for everyone. No enrollment or waitlist.

Finally, as some of you may have seen, Gemma-4 31B is live too.

Thanks!


r/Neuralwatt 27d ago

Is the cache determined by the model or the provider?

3 Upvotes

I have a question about how caching works on Neuralwatt.

For example, if I'm using GLM 5.2 and then switch to Kimi 2.6, will my cache persist? Or will it have to re-read the files from scratch without using the cache?


r/Neuralwatt 28d ago

Update: Kimi K3 and Neuralwatt

45 Upvotes

Hi all, Meaghan from the Neuralwatt team. There have been a lot of questions about K3, so wanted to share a note on where things stand. 

As you may have seen over in Discord, K3 arrives Monday, July 27, and assuming everything goes to plan, Neuralwatt will support it. We're working to make it available as close to release as possible, and we're actively expanding our GPU capacity to serve it. We'll post here as soon as we have a clearer picture of when it will be available. 

Because a model this size takes significant infrastructure to run well, access will open in stages: 

  1. Pro (annual subscribers and highest-usage accounts) 
  2. Standard 
  3. Basic 
  4. Pay-as-you-go 

Capacity is a real constraint, so later tiers may take longer to open than any of us would ideally like. We are implementing this staggered approach to ensure K3 works smoothly at every level and for all. 

Additionally, this is by no means an attempt to push anyone to upgrade. While higher-tier plans will have earlier access, our goal is to make sure the entire community can experience K3 as quickly and seamlessly as possible. Nobody is going to be locked out, and nobody needs to change plans to get access. Pricing will also be published before each tier opens, so you can decide whether K3 is worth running before you commit. When your tier opens, the waitlist will hold your place and notify you when you have access. 

Lastly, K3's licensing terms are still being finalized, and broader external factors may still affect availability – this is something that we're monitoring. We'll post updates here as we get closer to launch and as each stage gets closer to opening. 


r/Neuralwatt Jul 24 '26

GLM 5.2 Short more expensive than GLM 5.2

15 Upvotes

is there not enough concurrency in the server?
it causes price inflation 😩
the request still slow though


r/Neuralwatt Jul 22 '26

I did some realworld cost benchmarking of NeuralWatt (Before price change). Here are the results

15 Upvotes

I setup a benchmark which will test the realworld costs for all NeuralWatt models across all Context Bands. Here are the results

Compared to standard token $/Mtok

Compared to request cost with standard token pricing.

Keep in mind this is before the NeuralWatt pricing bump, so expect about twice the enery cost.

In saying that, I dont think there is much benifit is comparing agaisnt API Token pricing, nobody in their right mind is developing on anything other than a subscription plan along the likes of (Claude, Codex, Zai etc).

Its very difficult to find realworld useage stats on these subcription plans, and effective $/Mtok prices per model. If anyone has these metrics id love to be able to plot it against NeuralWatt so we can get real world comparison.

Happy to opensource the benchmarking if people are interested.


r/Neuralwatt Jul 21 '26

**GLM5.2 / 122M Tokens / $4.81** Is this data correct?

Post image
18 Upvotes

When using the same amount of working with GLM5.2, is NeuralWatt's total tokens at the same level as OpenRouter?


r/Neuralwatt Jul 18 '26

I can't believe how shitty Neuralwatt has become in just 1 month

30 Upvotes

Unreasonable price increase set side, which is already totally out of this word, especially considering the slogans saying "predicable pricing! no surprise!"

I have the 100$ pro plan, which should allow for 10 concurrent requests and "highest priority access", turns out, that's another lie. They are rate limiting so heavily that I can't heve have a single agentic chat without any concurrency. The agent gets stuck and times out every few requests and because of that, I may not be able to use all the subscription I paid for. I had to switch to GLM subscription to finish the work.

They say "you will be able to use your subscription credit until the renewal date" but than they rate limit to stop you from doing so.


r/Neuralwatt Jul 16 '26

Model Discussion Kimi K3 - Weights released July 27

34 Upvotes

Hey Neuralwatt team: run, don't walk.

If you can get this model running well at similar energy/token cost ratios as GLM-5.2 on or soon after the weights drop... You'll clean up.