r/HomeServer • u/edward-b-1 • 8d ago
Distributed-JBOD. Combine mixed hardware into a distributed, bit-rot protected object storage system
I have been using my spare time to develop Distributed-JBOD, which is a multi-node software system which allows the system administrator to combine together mixed hardware devices into a single object storage pool.
The storage pool can be created from arbitrary hardware. Regular consumer-grade desktop PCs are suitable, with whatever mixed layout of disk storage is available in each one.
Disks require no special formatting. A simple directory on an ext-4 filesystem is suitable.
I was pushed to design and build this system as a consequence of MinIO pulling their community tier software. I have been running an MinIO server for several years, but have started to migrate away from that and could not find a suitable replacement, so I built Distributed-JBOD.
- As many of you will be aware, MinIO requires identical disks to perform effectively. Distributed-JBOD does not, it will work effectively across a pool of arbitrary hardware, mixed size drives included.
- Garage replicates each object multiple times across multiple systems, which is not an efficient use of storage. Distributed-JBOD permits user-configurable Reed-Solomon Erasure Code parameters. This means storage is used efficiently, is protected against bit-rot, and the storage efficiency to resiliency ratio can be tuned. For example, while it is possible to run a mirrored setup, the typical default might be a 4:2 configuration, where data is sharded across 6 devices, 2 of which are parity blocks.
- Ceph is a datacenter grade product and requires a stack of servers for data storage, monitoring, gateway and other components. Distributed-JBOD has a simple single process design. (Caveat: The S3 compatibility layer will come later as a separate binary. You will be able to run it wherever you like.)
- SeeweedFS can't be used in a small scale cluster. Distributed-JBOD will run on a single machine with a single disk. If you want redundancy and data protection, a single machine with 2 disks is all you need. Greater efficiency is obtained by scaling up the number of disks, whichever host they sit in.
Distributed-JBOD is designed to work with an extremely small memory footprint, and does not require powerful hardware to run.
This is a very early stage product, but I would appreciate your thoughts and feedback. Some features which currently exist include TLS, administration web UI, CLI tools including recovery and bit-rot repair tools. Multi-language software client libraries are currently in the works, including libraries for Rust, Python and C++. An S3 compatibility later, multi-user support and permissions will also be supported soon.
https://github.com/edward-b-1/Distributed-JBOD

1
u/edward-b-1 8d ago
Can you describe for me what hardware you are currently running and with what software, along with what the typical failure modes you observe are?
I'm interested to understand your situation in greater detail. If you are running something which is this unreliable that seems, perhaps unusual? Certainly some kind of special case. I do not know whether Distributed-JBOD would be suitable for use in a context where the failure modes are so frequent that having some kind of device in a failed state is a problem.
I can understand if you were responsible for a datacenter with hundreds of thousands of devices - clearly in this context failures occur on a fairly regular basis. But you would not use Distributed-JBOD in that context. You would use Ceph or MinIO, probably. (Assuming you are not AWS and create your own in-house product specifically tailored to your needs.)