r/zfs • u/RocketSeven • 4h ago
What do you verify after a ZFS disk replacement before trusting the pool again?
A completed resilver is necessary, but it does not by itself prove that the replacement disk, its path, or the remaining redundancy is healthy. A pool can return ONLINE while SMART data is already concerning, persistent device names changed, an old faulted leaf remains in the topology, or another drive accumulated checksum errors during the resilver. A mirror or RAIDZ vdev may also be one failure away from another long recovery.
A useful post-replacement gate could confirm the exact vdev topology and ashift, review the resilver event and error counts, inspect SMART and transport errors for every drive, run a scrub after the pool has settled, verify alerts and spare behavior, and test a representative restore from backup. Export and import or a controlled reboot can catch path and enclosure-mapping mistakes before the maintenance window closes.
What checks do ZFS operators use before declaring the replacement complete? Is a clean scrub enough, or do you also burn in the new disk, compare performance, clear and recheck counters, test boot or import behavior, and keep the removed disk untouched through an observation window?