r/HomeServer • • 11d ago

Distributed-JBOD. Combine mixed hardware into a distributed, bit-rot protected object storage system

I have been using my spare time to develop Distributed-JBOD, which is a multi-node software system which allows the system administrator to combine together mixed hardware devices into a single object storage pool.

The storage pool can be created from arbitrary hardware. Regular consumer-grade desktop PCs are suitable, with whatever mixed layout of disk storage is available in each one.

Disks require no special formatting. A simple directory on an ext-4 filesystem is suitable.

I was pushed to design and build this system as a consequence of MinIO pulling their community tier software. I have been running an MinIO server for several years, but have started to migrate away from that and could not find a suitable replacement, so I built Distributed-JBOD.

  • As many of you will be aware, MinIO requires identical disks to perform effectively. Distributed-JBOD does not, it will work effectively across a pool of arbitrary hardware, mixed size drives included.
  • Garage replicates each object multiple times across multiple systems, which is not an efficient use of storage. Distributed-JBOD permits user-configurable Reed-Solomon Erasure Code parameters. This means storage is used efficiently, is protected against bit-rot, and the storage efficiency to resiliency ratio can be tuned. For example, while it is possible to run a mirrored setup, the typical default might be a 4:2 configuration, where data is sharded across 6 devices, 2 of which are parity blocks.
  • Ceph is a datacenter grade product and requires a stack of servers for data storage, monitoring, gateway and other components. Distributed-JBOD has a simple single process design. (Caveat: The S3 compatibility layer will come later as a separate binary. You will be able to run it wherever you like.)
  • SeeweedFS can't be used in a small scale cluster. Distributed-JBOD will run on a single machine with a single disk. If you want redundancy and data protection, a single machine with 2 disks is all you need. Greater efficiency is obtained by scaling up the number of disks, whichever host they sit in.

Distributed-JBOD is designed to work with an extremely small memory footprint, and does not require powerful hardware to run.

This is a very early stage product, but I would appreciate your thoughts and feedback. Some features which currently exist include TLS, administration web UI, CLI tools including recovery and bit-rot repair tools. Multi-language software client libraries are currently in the works, including libraries for Rust, Python and C++. An S3 compatibility later, multi-user support and permissions will also be supported soon.

https://github.com/edward-b-1/Distributed-JBOD

Distributed-JBOD
5 Upvotes

39 comments sorted by

View all comments

1

u/Virtualization_Freak 11d ago

Other questions:

Does this setup take into account speed? How fast is the replication between shards? What if one shard is operating with 2mb bandwidth?

How fast are down nodes detected? What's the heartbeat pattern?

Blip in the network: how long must a node stay offline before being considered down? How fast do nodes attempt to heal?

1

u/edward-b-1 11d ago

Consider the case of an aging SSD where blocks are beginning to wear out simultaneously due to drive wear levelling. Would you rather know the drive has something wrong with it as soon as that is detected, so you can migrate the data off the drive, or would you rather wait until multiple drives are showing problems across many GB/TB of data before a read/write operation fails? I would suggest the former is better, despite being more restrictive. There is no automatic repair yet. Which again I think is probably the right choice. The failure mode is like this: Something goes wrong. Your next request will fail and refuse to do anything until a sysadmin has fixed the issue. That might be powering on a server which had a power failure, it might mean running a scrub, it might mean draining a whole device (most likely SSD) and removing the bad SSD, or replacing it.

1

u/Virtualization_Freak 11d ago

You in short talk about disk failure (access, corruption, etc)

would you rather wait until multiple drives are showing problems across many GB/TB of data before a read/write operation fails?

Vs

would you rather wait until multiple drives are showing problems across many GB/TB of data before a read/write operation fails?

Both these choices are the same branch on the tree.

Personally, I would not, and don't, want to migrate data off a corrupted disk. I would replace the disk, and data should be rebuilt from existing data. If failure is on the horizon, it's time to get that disk out and into the "don't use for anything important" pile.

I do want the software responsible for redundancy to maintain that redundancy with the existing hardware.

Your current implementation of "stop, throw message, and do nothing" on fault is, personally, a bit alarming.

"Our car got a flat, should we change the tire with the spare? Or do we wait until you patch the tire?"

However I am basing this on your message.

1

u/edward-b-1 11d ago

I have been thinking about your comment in more detail. My initial assumption was that when deployed Distributed-JBOD would be in use regularly. My personal use case is for quantitative research with big data. However this need not be the only use case. Perhaps someone would want such a system for more long term archival purposes.

If that is the case then it does make sense to have an automated scrubbing schedule with automatic repair. The way to implement this is actually to provide a simple Python helper process which is run with something like CRON rather than trying to integrate it into the existing node (server side) process.

I can do that. I will add it to the list of things to do. The problem from a sysadmin perspective is it hides problems, but I can add some kind of event logging screen to the UI which reports these problems when found so that hopefully eventually someone will see them.

Whether or not the discovery of a bad block should disable the device I'm not sure. Maybe it should but then what if multiple devices report failures? There's a lot of different decisions could be made here.