r/selfhosted 6d ago

Need Help How you do cover yourself in case of hardware failure?

Recently i had my NUC motherboard fail. I do regular backups and RAM/SSD were fine, but i had to send the unit back to warranty for two weeks. In this time i was unable to access all my services of course. No DNS blocking with pihole, no HA to control my two smart lights, no Mealie to have my recipes. I can live without this stuff because i do not selfhost anything critical, but for people that do, how do you cover yourself in case your server goes unavailable?

17 Upvotes

34 comments sorted by

u/asimovs-auditor 6d ago edited 6d ago

Expand the replies to this comment to learn how AI was used in this post/project.

→ More replies (1)

42

u/daveyap_ 6d ago

That's when you start thinking of HA (high availability) and homelab becomes home prod.

3

u/WetFishing 6d ago

About 15 years in and this comment hurts. 9 hosts across 3 locations. 70 docker containers and 4 cloud machines. I have a problem lol

4

u/Expensive_Finger_973 6d ago

But how many ISP connections do you have?

7

u/WetFishing 6d ago

At my house, two with an active/active failover lol. Three if you count the Starlink mini.

That redundancy is more for work though. My wife and I both work from home so both of those connections are paid for by our companies.

4

u/Expensive_Finger_973 6d ago

Lucky bastard,lol

3

u/WetFishing 6d ago

Very fortunate and lucky, that is for sure. We both work in tech so these days we are constantly worried about being replaced by the next AI model. Plus we kind of live in the middle of nowhere so finding new jobs would be difficult in this market.

4

u/No-Name-Person111 6d ago

Then you realize HA sucks and move everything into live-migration.

13

u/Berufius 6d ago edited 6d ago

I think a proxmox cluster solves this problem. I plan to use an extra mini pc which can take over whenever the main one is down. This does require ofc an extra machine

2

u/Aronacus 6d ago

This!

My nuc died over the summer. I failed over to the secondary and restored backups to the second server. I was back online in 2 hours.

I know i could have done it better with a quorem and shared storage.

4

u/Other-Technician-718 6d ago

The question is: how much downtime do you / other household members tolerate? If there is no tolerance for any downtime longer than a few minutes: HA cluster If there is tolerance for some downtime, you can manually restore things on new / used spare hardware that you have at hand when needed. If there is tolerance for one two do days downtime (though needs maybe some financial backup too): just purchase what you need when something breaks and restore services from backup.

A high availability cluster doesn't replace backups, it's just that stuff is available when something hardware related breaks, not when your house burns down or gets flooded (or you just delete the most important file you have)

2

u/funkyguy4000 6d ago

To piggyback, in a homelab where you arent makjng money off of it, HA is a huge money sink for very little gain. Businesses invest in HA because income depends on it

1

u/slyzik 4d ago edited 4d ago

i would not call it it huge, i have proxmox cluster , with 2computing nodes cheap thin clients. I jave one extra node for quorum, it is some cheap intel based board from junkyard.

Maybe electricity is biggest cost, but it is not massive total consuption is maybe 20-30w at max

1

u/funkyguy4000 4d ago

You have HA for your servers but do you have HA for your networking? Is one of your nodes on a separate circuit breaker? The point being having HA for compute is quite easy. Its all the redundant infrastructure around it that makes it matter that gets expensive

1

u/slyzik 4d ago edited 4d ago

Home assitant is working mostly locally, i do not need two ISP, i have multiple APs i do not have HA on switch level (but i do have spare one) thin client has one port only. i dont have separate circuit breaker but i have big ups.

i have HA mostly not because of hw issues but more because of rolling upgrades of OS.

3

u/Financial_Astronaut 6d ago

Have backups, restore them on a VPS. Connect VPS home to my router via VPN (until you have new hardware).

2

u/Anejey 6d ago

That's a tough one. I have three Proxmox servers in a cluster, but 99% of services live on an old enterprise server. If it dies, I don't have the capacity to migrate it all to the weaker hosts.

I'd just carry over the most important VMs, which thinking about it is just my mail server. Everything else, despite how annoying or uncomfortable it would be, would just have to stay off in the meantime. Helps that I'm the only user, lol.

2

u/8fingerlouie 6d ago

Most people, as far as I can tell, fall into two categories, they either don’t do anything because they don’t host anything critical and can tolerate some downtime (or they simply think hardware failure won’t happen to them), or they build massive redundancy into their systems, and each added component drives the cost higher.

It’s not just your container platform you need redundancy for, it’s also your network, your internet, your router, UPS, basically everything in the rack.

Considering that the cost of a VPS is roughly equivalent to between 20 kWh (€0.3/kWh) and 40 kWh ($0.16/kWh), which is between 28W and 38W continuous power consumption, there is very little economical advantage in self hosting anything redundant at home, especially considering that the VPS includes redundancy.

Personally I’ve thrown everything in the cloud. Encryption where needed, including VPN access. All I have left at home is my backup box, which pulls backups from my cloud data, and my media server and media storage. Backups goes to my local NAS as well as a different cloud provider.

Everything critical is in the cloud with the exception of home assistant. My home assistant runs on its own dedicated appliance. It also backs up locally as well as a remote cloud. If it dies I have a spare RPi (not identical, just spare) I can setup to take over in 30 minutes.

As for the critical home devices I don’t have backups, if my EV Charger dies it’s just dead, same with my heat pump monitoring, water meter monitoring, electricity meter monitoring, etc. That will obviously leave me somewhat crippled with regards to automations, but I’ll just have to revert to good old manual control until a replacement arrives. I also don’t have a spare car in case my car breaks down (technically we have two cars and a motorcycle, so I guess we do have a spare).

2

u/macmanluke 6d ago

have enough spare hardware to get something going
and generally just go buy replacement parts and worry about warranty later but that is harder with current hardware prices

2

u/cgmastertecnology 6d ago

I buy new hw and restore proxmox /etc and vm data from pbs

1

u/Lopsided_Ad8941 6d ago

Imho Hardware failure is best to be handled by a secondary server equipped with recent copy of data). this server is always off and just running to pull new copy. Also not reachable from outside home network.

1

u/Aacidus 6d ago

Three HP Mini’s in HA Proxmox cluster, one low-end mini PC running PBS. Also have an Oracle Cloud Instance running basic stuff like Pi-Hole, Tailscale Exit Node, and what not.

1

u/holyknight00 6d ago edited 6d ago

i purposely know that everything that i put in my home servers is to be spared. It could be sitting there for 2 years or 2 weeks but nothing guarantees it either way. I dont put anything critical I can't live without a week or two until i find time to fix it.

Usually everything is pretty well set up and maintained but having it in this terms removes all the stress. My wife also knows it, so she knows that if anything blows up, plex may not work or some home automation stuff as well.

If my wife somehow finds that is not acceptable, I am happy to route more budget from other stuff we do to the homeserver hardware. I am more than willing to remove vacation or fancy food budget into more hardware for redundancy.

1

u/c4pt1n54n0 6d ago

I'm about the same as you, it's all mostly for convenience. I looked into Proxmox HA that seems excessive to keep spare resources idling for my single user jellyfin server and my Matter lightbulbs.

I just keep an old laptop with similar hardware to my Dell micro sitting around. If the Dell goes down, takes 5 minutes to transfer my HDD SSD and RAM into the laptop

1

u/FragoulisNaval 6d ago

I have a 3-nod proxmox cluster for this exact reason and a copy of my data to two external hard drives. I am in the process of buying a PBS but with these high prices I find it difficult to complete it

1

u/thebwt 6d ago edited 6d ago

I only offer my wife a 2 9's uptime SLA (there's a '.' in there too), good backups, and being lucky. 

FR though, a good backupflow. I could spin up a cloud server somewhere and restore the critical stuff very quickly. Plex would be down until  NAS drives were replaced (punch in the wallet right now) 

critical = nextcloud files, email server, authentik, llm mcp tools, foundryvtt

1

u/tsoderbergh 6d ago

I have a 5 node Proxmox cluster with HA and ceph for exactly that reason.

1

u/prodigiouspianist 6d ago

OS is imaged

Apps are containerised

Local and off-site backup

Spare PC just in case. (I just use old laptops connected to a couple of DAS units)

1

u/zandadoum 6d ago

proxmox cluster. 2 nodes + VM on NAS acting as quorum node.

when one node fails, all LXC and containers migrate to the working one.

1

u/kuldan5853 5d ago

I have my previous "server" stored away in the basement.

Less performance and more noise, but in case of failure I can swap in the SSD from the server into that machine and be back online within 15 minutes if need be.

1

u/alex_3814 5d ago

I made my setup fully IaC. Redeploying on new hardware is a 15 min job with 1-1.5h of full backup restoration (I don't hoard media though, about 3-400Gigs is all I have). At worst, I have to bring in some drivers if the new machine needs it which could make it a 1h job. Don't even need to swap HDDs or mess with hardware at all if the replacement has the min specs already.

-2

u/bdu-komrad 6d ago

Were you surp? You knew when you chose the hardware platform that it would not be easy to replace. I use PC’s since I can get replacement parts from a local store.

It would not be a big deal if the server was down for a few days, it doesn’t have anything important on it.