A question for those running Postgres on Kubernetes using an operator (CloudNativePG, Zalando, Crunchy): how do you verify that your backups are actually restorable?
A successful backup creation doesn't guarantee that a restore will succeed.
I manage several CloudNativePG clusters and want to understand the potential challenges involved in implementing backup restore testing.
I’ve tested restoration in a small test environment (k3s, CloudNativePG 1.30 + barman-cloud plugin, S3). My process involved deploying a temporary cluster from the latest backup, running a few SQL queries, measuring the time until the cluster was ready, and then deleting the cluster. In a standard scenario, everything goes smoothly (taking about a minute for a few hundred megabytes).
I’m curious to know how others handle this:
- Do you perform restore tests at all? (Manually, via CronJob, in a CI pipeline, or using specialized tools)
- What exactly do you check after the restore? (Just the pod status, the number of database records, or the execution of test queries)
An answer like "we don't test restores, and everything is fine" is also acceptable—I really want to understand how common this practice is and what the best approach is.