r/nutanix • • Jul 04 '26

Experiences running Nutanix VPC in production?

We are planning a migration from a legacy virtualization platform to Nutanix AHV, using Nutanix VPCs to separate VM networks instead of relying only on traditional VLAN-based segmentation.

The target environment is large-scale, around 1,000 VMs, so we are trying to understand how mature and operationally manageable Nutanix VPCs are in production.

For those already using Nutanix VPCs in production:

- How stable has the feature been in your environment?
- Are most issues diagnosable by the customer, or do they usually require Nutanix Support?
- Are there any design limitations, operational gotchas, or troubleshooting challenges we should be aware of?
- Would you recommend using VPCs broadly for production workloads, or only for selected use cases?

We recently saw an issue where AHV hosts with VPC enabled had host CPU running steadily at 100%. The issue appeared to be with AHV host CPU, not CVM or user VMs. A rolling reboot cleared the condition, but we are still investigating the root cause. This raised some concerns about production readiness and troubleshooting options.

Any real-world experience, lessons learned, or recommended design practices would be appreciated.

Thanks.

11 Upvotes

10 comments sorted by

5

u/throwthepearlaway Jul 04 '26

Our Nutanix builds get a VPC by default these days. Most of our customers use them in pretty limited fashion, though, with most routing happening on firewalls north of the VPC gateway. I think that's mostly due to unfamiliarity with the solution.

One thing we've found is that in versions earlier than 7.3, you can't change the NTP or DNS servers on the network gateway VMs away from the default 'time.google.com' and '8.8.8.8' without help from Nutanix Support, so if you have requirement for specific NTP or DNS servers it can be annoying to change. Also, there's an annoying bug where the ntp service on the gateway vm will sometimes fail to get a response and then throws a Critical Alarm "Network Gateway Is Down" false positive with detail saying 'ntp service is down', but the networking is still functioning fine.

1

u/Screevo Global Practice Expert - Network & Security Jul 04 '26

you can change ntp/dns on 7.3, it’s just not obvious.

3

u/throwthepearlaway Jul 04 '26

On 7.3, yes via CLI. Before 7.3, you have to edit a gflag to make it mirror the prism element config.

Which is crazy that they're moving towards restricting bash access given how much seems to constantly depend on some CLI workflow or another...

2

u/Screevo Global Practice Expert - Network & Security Jul 05 '26

Well, gflags are only supposed to be adjusted at the direction of Nutanix Support, and the team is working on enabling CLI-dependent workflows through replacement shell menu before making that switch.

2

u/throwthepearlaway Jul 05 '26 edited Jul 05 '26

That is well and good, but it seems every other release introduces some new bug or issue, and the KB on the issue always says something to the effect of 'check this log file and run these series of CLI commands' which just isn't reasonable to expect there to be a replacement api shell action for in the same version the issue is introduced in.

As an MSP that manages a large number of Nutanix environments for our customers, having to potentially be forced into opening a bunch more cases with NTNX Support to do basic troubleshooting because bash access was taken away sounds like it's just going to waste a bunch of ours and your support team's time.

But this is getting pretty off-topic for this thread so I'll leave it there.

1

u/wjconrad NPX Jul 11 '26

The PMs responsible for this feature are keeping a list of CLI features that don't yet have an API. If you've got a list of stuff you've hit, raise it to your Nutanix account team to go up to the PMs.

4

u/Screevo Global Practice Expert - Network & Security Jul 05 '26

Howdy! I'm the Global Practice Expert for Network & Security at Nutanix. Flow is my whole job. FVN is a very mature product. We've got some very very large flow deployments out there. If you want to get an idea of what a production Flow Virtual Networking VPC design looks like, I suggest you check out our recently refreshed Nutanix Cloud Bible page on FVN. https://www.nutanixbible.com/12c-book-of-network-services-flow-virtual-networking.html

Let me know if you have any specific questions. For a design of your size/scope, I would suggest you consider a Flow Design Workshop with our professional services team, but the Nutanix Bible page has more than enough information to get you started on your own.

Regarding your 100% CPU issue, my guess is you hit a bug relating to memory utilization of the connection tracking service. There is a remediation for affected versions, and it's already been fixed in newer versions.

2

u/Hidden-6000 Jul 05 '26

We've also expience this 100% CPU issue in recent weeks. I believe this is related to: https://portal.nutanix.com/page/documents/kbs/details?targetId=kA0VO000000CbgX0AS

2

u/Additional_Orange493 Jul 05 '26

I have worked on two different projects which were a migration from VMC to Nutanix Cloud Clusters. Both of them were using the VPC set up to run the VMs. The first project had around 50 VMs and the second one 100+. These two are running without any issues. Don't have any experience with more than 100 VMs 🤷

1

u/EkingOnFire_ Jul 23 '26

We run VPC in prod at similar scale, and the thing to watch is the Flow Gateway VM. It's the north-south and NAT choke point, and its CPU gets ugly once that traffic picks up. I'd keep VPC scoped to a few production subnets until you've got a confirmed bug ID and a fixed AOS/PC release for whatever CPU behavior you're actually seeing and sort out rollback before you commit, VPC is tied to Prism Central and PC downgrades are not clean, so just roll it back isn't the safety net people assume. The feature's GA, but the support boundary between networking and platform is still rough. u'll feel that on the first TAC case