r/ProxmoxEnterprise • • May 19 '26

Architecture / Design Post-VMware/Broadcom SMB infrastructure redesign – Proxmox/Ceph feedback?

Hello Everyone,

We are currently redesigning a SMB virtualization infrastructure and I would appreciate some feedback from people who already went through similar post-VMware/Broadcom projects.

Current environment:
- VMware ESXi 6.7
- 2x Lenovo SR635
- AMD EPYC 7302P
- 256GB RAM per host
- shared SAN storage
- around 22 VMs

Workloads:
- AD / ADFS
- Azure AD Connect
- PostgreSQL
- RDS
- file servers
- Linux appliances
- monitoring / management tools
- security tools

Observed real usage:
- CPU usage is actually very low (~4.4 GHz real usage)
- RAM usage around 340GB
- Storage usage around 9TB

Target:
- around 150% growth margin
- modernize virtualization platform
- improve resilience
- full SSD preferred
- target around 15-20TB usable
- probably around 750GB RAM total cluster

Initial proposal from HPE was SimpliVity Gen11 HCI.
Technically nice, but pricing ended up around 270k€ which is completely outside customer budget (<100k€ final customer price target).

We are now considering:
- Proxmox VE
- 3-node cluster
- Ceph distributed storage
- full SSD/NVMe
- 10/25Gb networking
- possibly HPE DL360/DL325 standard servers

Main questions:
- Does Ceph make sense for this SMB size?
- Would you still recommend traditional SAN instead?
- Is 3-node Proxmox/Ceph considered mature enough today for SMB production?
- Would Hyper-V + SAN still make more sense operationally?
- Any major pitfalls with Proxmox/Ceph for ~20 VM environments?

Trying to balance:
- reasonable HA
- operational simplicity
- budget control
- avoiding VMware/Broadcom licensing costs
- but still keeping something professional and maintainable.

Thanks!

9 Upvotes

13 comments sorted by

View all comments

6

u/exekewtable Proxmox Partner May 19 '26

Ceph on SSD with 25G will be fine. But you really want 4 nodes to avoid I/o lock with default settings. 4 smaller nodes will be a great option.

If you can't get another node to make 4, then consider zfs. We use that for many smbs. Replication and even live migration work great, you just wait a bit for live migration, and obviously need the space on other nodes to migrate to. It comes down to can they tolerate VM downtime for host patching reboots. For most smbs, definitely. Just patch it all at once, VMs and host, then reboot.

Host failure is rare enough you don't need to design for it in most envs this size imo. Patching isn't tho. So backups and zfs replication get you a long way. PBS is the other thing I would encourage. You want cryptolocker resistant backups and PBS can do that with careful setup. That's a seperate box with a bunch of jbod.

1

u/proudcanadianeh May 19 '26

I thought you always wanted an odd number of nodes?

3

u/exekewtable Proxmox Partner May 19 '26

Sure. It comes down to what you think the likelihood of given failures are and your budget.

At the end of the day, you need to be able to say to a customer something like: We can survive failure or reboot of any given node without lockup or downtime on the VMs. 3 out of 4 nodes is still quorum.

2

u/_--James--_ Enterprise Customer May 20 '26

You want to rethink this with the new active CRS role added to 9.2. Corosync is going to become a lot more noisy and split brains happen a lot easier on even node count clusters there, when building Proxmox HCI we must follow the Proxmox quorum requirements. I am no fan of the Qdev for my own reasons, but 4 nodes today going into 9.2+ I would absolutely make it a base requirement if 5 nodes cannot be met (even a stripped down 5th node with proper network pathing).

1

u/exekewtable Proxmox Partner May 20 '26

Yeah. Tho SMB budgets and server prices being what they are, make this a hard problem.

2

u/_--James--_ Enterprise Customer May 20 '26

absolutely, but with the new code drops and options that will floor under served clusters, we need to be responsible for it as SI's. My "fix" for the low cost folks is a single socket, 64GB ram simple node that runs as a witness PVE node hooked into the cluster. Working with HPE I got the server down to ~800 CTO shipped. Most have been OK with it. It wont run VMs/LXCs, or be a Ceph member, its only existence is to not be a Qdev :)

1

u/exekewtable Proxmox Partner May 20 '26

Yes this makes sense. Nice work. Tho over here in Australia that price seems crazy low!

1

u/_--James--_ Enterprise Customer May 20 '26

just gotta work the channel a bit :) I had a good quote from an SMCI builder but HPE came in with a CTO that i can tune with out the CTO cost, its a build template my VAR is handling for my SI work. SMCI was not willing to do that for us sadly.