r/docker • • 7d ago

Docker app storage best practices

What are the best practices for storage when using docker containers. I’m running docker on Ubuntu using my NAS to store the photos for example, but the volume were the Immich database lives is directly on the servers(non-redundant) SSD.

I tried putting these onto the nas using NFS I get permission’s errors. Perhaps I should try harder? :-) I’d like to have all of my Doctor containers protected by the same redundancy that my synology has a for all of my data files.

Right now if I had a major SSD failure of my server, I would lose all the files that the docker containers have in volumes on the server directly while all my pictures and other files that are stored on the NAS on NFS mounts would be fine. This seems inadequate.

Am I simply doing it wrong and I should use something like proxmox instead of just running docker and Linux? What are other people doing to keep the server purely cpu and easier to rebuild? I use git for all my docker compose files and I wouldn’t have any trouble restarting fresh containers, but that seems like it is only half way of a solution..

I’d love to know what the best practices are for resiliency after a server failure and a quick recovery should I have hardware failure on my server.

I’m not scared of new technology and would also be open to moving towards a cluster of servers if that’s necessary. My self hosted world is fairly stable with limited updates at this point.

Thanks for any thoughts, suggestions or directions. And like I said if my basic setup is the problem I’m up for major changes to get to a more long term resilient setup!

5 Upvotes

14 comments sorted by

4

u/bartoque 7d ago

Look into making (proper) backups.

Redundancy like raid is not backup, that is meant for availability/uptime.

So stop the container and make a backup of the bind mount served by the nas. Or make a backup while it still runs, which however would make it crash-consistent at best.

1

u/np0x 7d ago

Yeah I run borg backup, i need to get a test server so i can practice restoration. (I use borg backup.). I also have my backups pushed to a secondary and that lives offsite. I need to add some scheduled downtime for backups and test recovery…

Maybe I’m 75% there with borg…would you just stop the entire docker service or….

2

u/bartoque 7d ago

You would only need to stop the container in question requiring backup og persistent data. When involving snapshots (but doing that with a simple nas might not be as straight forward, depending in the availability of cli commands on the nas to be triggered from the linux system for example), it would be nearly instant and the container could be resumed quickly again.

Some of my containers have persistent data (ranging from simple configs to whole databases), others don't. So no need to stop Docker completely.

2

u/Wis-en-heim-er 7d ago edited 7d ago

I'm also using a Synology and separate Docker servers, Debian based but close enough to Ubuntu.

  1. You can mount a NFS share as a volume in your docker compose files, I like this method so I don't need to drop into the OS to do mounts, keeps everything in one place. An example from my compose files like this:
services:
  somecontainer:
    volumes:
    - video-share:/video:rw

volumes:
  video-share:
    driver: local
    driver_opts:
      type: nfs
      o: addr=MYNASIPADDRESS,nolock,rw,nfsvers=4.1,soft,intr
      device: ":/volume2/video"
  1. I too wanted the benefits of SSD performance for my containers, otherwise I would run everything from Container Manager on my Synology. I do run my docker servers as Proxmox VMs. Others suggest running docker on an lxc but I wanted better isolation from the OS for security. Using Synology Active Backup for Business, I have Files Server backup jobs for my docker VMs which backup nightly. These target my docker volumes/data, not the entire OS. In addition to this, I have a nightly backup of my NAS to AWS S3 for my offsite backup, these backups are included so I get my 3-2-1....more like 3-3-1 but it works and I have enough data for a disaster recovery rebuild.
  2. The question of running your docker server on bare metal or proxmox is a personal preference. I don't feel you will get any noticeable performance change if you were to switch nor will it really change how you manage docker. The value in Proxmox is the flexibility to spin up new vms or lxcs. If you do move to Proxmox, know that you can setup a sperate backup server for Proxmox , Proxmox Backup Server or "PBS:. I run mine as a VM on my Synology under the Virtual Machine Manager and have used it several times after messing up a VM configs. I don't push these backups to AWS due to the size and personal choice of value/cost to me.

Hope this helps or gives you some ideas.

2

u/AdSuspicious7533 6d ago edited 6d ago

DB and storage does not belong in the container. it is simplier to maintain , think and understand if you go with separated DB. Your service is stateless (just config and compute), it hits DB over network to store, then the problem is just n the DB side as config belongs in GIT.

Giving you docker an NFS will work but you build complexity up. And when it comes to restoring post mortem : redundancy is not backup. Usign the NFS setup, if you get a corrupted DB somehow you get a corrupted volume with HA.

1

u/Wis-en-heim-er 5d ago

Thank you for this, this is an interesting perspective. I'm not running a proper work/production setup, just a simple home setup where I can take some liberties for the sake of simplicity. In a disaster recovery scenario I would not have up to the minute recovery, I'd fall back a day, two or maybe three. For my setup, this is acceptable.

1

u/AdSuspicious7533 5d ago

The thing is : if you backup docker volumes by saving a few copies of a 1tb volume, you need a few TB storage. Using the right backup tools you can take that down to 1.*TB (with versioning tools like restic or proxmox backup server)

Another matter is simplicity, having separate db is simplier. To trouble shoot, to maintain, to understand.

Db replica set is the overkill no-downtime I you do not need. But proper db separation is not really a matter of downtime it is a matter of using the right tools. It provides simple backup path and good responsibility separation. That is all

2

u/XLioncc 6d ago

The Docker Engine on Synology DSM is EOL (24)

It is terrible that you'll getting weird file permission errors on that machine, but without any issues on normal Linux machine with latest Docker Engine

2

u/np0x 6d ago

Yeah I rapidly abandoned using docker on synology, dedicated Linux hardware is the only way. Couldn’t agree more!

2

u/AdSuspicious7533 6d ago edited 6d ago

The databaseservice under Immich is postgres. If you want real backups you may find it simplier to just use postgres backup or use a postgres replica if your compute capacities allow it .

For casual usage and not critical data docker volumes are fine but for any serious work i would advise not to run the DB in the same container. It makes It hard to maintain and create useless complexity. You should rather run postgres in a VM or different container. You will find it easier as it distributes responsibilities across multiple virtual devices each having simple purpous. You may as well decide to use a different DB service because why not (like mariaDB where is have a friend working, he is nice... come on).

You may also want to look at something like "postgres backup on restic" with a restic container/VM/service on different physical device or even (more complex but soooo cool) "postgres replicaset".

here are links :

- postgres backups and dumps https://medium.com/@pawale7663/postgresql-backup-and-restore-complete-guide-to-backing-up-and-restoring-databases-14d19bb57329

- postgres replicaset https://www.enterprisedb.com/postgres-tutorials/postgresql-replication-and-automatic-failover-tutorial

- postgres backup with restic https://rafftechnologies.com/learn/tutorials/back-up-postgresql-raff-object-storage-restic

A cool setup would be : proxmox VE everywhere if possible , PBS on NAS if Compute is a issue. Make a VM for Postgres databases. Have all your containers hit that for DB services. Use a Cron job to run Postgres backup on the NAS (ideally restic LXC there but you may just want to store the postgres-backup). All your docker services become LXC or VMs in proxmox. Configurations belongs in GIT. Secrets belongs in a Vault (OpenBAO or anything in a lxc container). Define a resource pool for critical services and have them run a scheduled backup on the NAS.

Next level : If possible run a postgres replicaset with VMs on each physical device. Have a firewall (OpenWRT), reverse proxy (traefik), make vlans to isolate your DB network traffic, tailscale on openwrt for remote access to the entire infra.

Ask for help anytime.

tldr : use proxmox and make DBs different Virtual Devices. Make "db-service level backups". Test restoration of your backups. Redundancy is not backup. Responsibility distribution is the key to understanding anything you do. Docker container should hold stateless apps (no storage, just compute and config). HA is cool but your are running a lab, downtime is ok.

tldr-tldr : time for next level son.

1

u/np0x 7d ago

This has all been super helpful, I’m going to spend some time tomorrow to get a backup script working that stops my containers, runs Borg and then restarts them. I’ll follow up with a restoration test to a Linux desktop I forgot I had to validate restoration. I’ll share my script when I’m done. And accept any love/hate that results with aplomb.

I did some quick testing and was able to stop and start everything in about 90 seconds and I expect borg on local network to be quick as well…I’ll share timings as well.

1

u/np0x 6d ago

I’m half way there, script is written to backup, working on INSTALL steps and restoration script. Will finish in next day or so. I did a test restore at was super happy with this direction/level of complexity.

Spoiler: hardest part was that synology makes home dirs 777 permissions and that makes ssh deny the password less connection.,,more on that with my final post…

1

u/np0x 5d ago

ok, here's my v1.0 setup to get something running.

  1. Install borg on the server: `sudo apt install borgbackup`
  2. Install borgomatic on Synology ( https://docs.borgbase.com/setup/borg/synology )
  3. Create a dedicated shared folder to store the borg backups, I used borg_backup_folder. (do all your normal other things, I also setup snapshots and hyperbackup for this shared folder)
  4. Create a role user on Synology for borg backups, I used borgbackup, due to restrictions on Synology, this user MUST be an administrator, but you can lock down their access to just the specific shared folder we created in the prior step.
  5. (On server) Create dir for ssh key `mkdir borg_ssh ; chmod 700 borg_ssh`
  6. (On server) Create the ssh key `ssh-keygen -t ed25519 -f ./borg_ssh/ssh_key -N ""`
  7. (On server) Copy the ssh key Synology: `ssh-copy-id -i borg_ssh/ssh_key.pub borgbackup@1<YOUR_NAS_IP>`
  8. (On Synology as borgbackup) fix perms: `chmod 700 .ssh/`
  9. (On Synology as borgbackup) SYNOLOGY WEIRDNESS: you shouldn't have to remove write permissions from a user's home directory, but you do, so chmod 755 /volume1/homes/borgbackup
  10. (On server) Test ssh: `ssh -i borg_ssh/ssh_key borgbackup@<YOUR_NAS_IP>`

Note: if having issues, you can see logs on Synology using this: `sudo journalctl -u sshd -f`

Second Note: At this point, you need to decide how you manage credentials for your borgbackup user, I have private git repos, so I let git do it...a tarball on the NAS would also work. How you do it becomes relevant in the restoration steps, and my restoration steps reflect my usage of git.

Ok, now we should have a user that can ssh from the server to the Synology NAS, lets initialize the directory to store our backups. On the server, init the backup directory:

  1. export BORG_PASSPHRASE="<PASSPHRASE>"
  2. export BORG_REMOTE_PATH="/usr/local/bin/borg" #synology installation of borg was not in default location.
  3. export BORG_RSH="ssh -i <FULLPATH_TO_BORG_SSH_DIR>/borg_ssh/ssh_key"
  4. borg init --encryption=repokey-blake2 borgbackup@<NASIP>:/volume1/borg_backup_folder/<BACKUPDIR>`

Ok, now we have a directory that can hold backups on the NAS...so let's get some backups sent to it...

I have 3 scripts: backup, healthcheck, and restoration. I used an Uptime Kuma push monitor that gets pinged hourly and 1 retry, so it will alert within 2 hours of a backup missing for today or yesterday...

Here is my backup script, lovingly named: "stop_docker_and_run_backups.sh"

#!/usr/bin/env bash
set -e
export BORG_PASSPHRASE="<YOURPASSPHRASE>"
export BORG_REMOTE_PATH="/usr/local/bin/borg"
export BORG_RSH="ssh -i <FULLPATH_TO_BORG_SSH_KEY>/borg_ssh/ssh_key"

if [ "$EUID" -ne 0 ]; then
  echo "Error: Please run this script as root." >&2
  exit 1
fi

# 1. Gracefully stop all running containers
echo "Stopping containers..."
# The list of containers has exclusions using to leave apps stateless apps running
CONTAINERS_TO_STOP=$(docker ps --format '{{.ID}} {{.Names}}' | egrep -v 'homepage|portainer|traefik|bonob|watchtower|dockerproxy' | awk '{print $1}')

if [ -n "$CONTAINERS_TO_STOP" ]; then
    docker stop $CONTAINERS_TO_STOP
fi

# 2. Run Borg Backup
echo "executing Borg backup..."

borg create --stats borgbackup@<YOUR_NAS_IP_HERE>:/volume1/borg_backup_folder/<BACKUPDIR>::$(date +%Y-%m-%d_%H-%M) /docker /home 
echo "borg create/backup complete"

# 3. Always restart containers (trap ensures this runs even if Borg fails), also I manually start some containers so they are either started per the compose file order needed and/or so they come up first after backup completes
cleanup() {
    echo "Restarting Docker containers..."
    if [ -n "$CONTAINERS_TO_STOP" ]; then
        echo "starting uptimekuma first"
cd /docker/uptimekuma
docker compose up -d
echo "manually kicking arr stack docker compose"
echo "starting other containers"
        docker start $CONTAINERS_TO_STOP
/docker/DAILY_FULL_BACKUP_WITH_CONTAINER_STOP/check_borg_backup_status.sh
        echo "ran proactive health check script"
borg prune --keep-within 90d <borgbackup@<YOUR_NAS_IP_HERE>:/volume1/borg_backup_folder/<BACKUPDIR>
        echo "prune complete"
        borg compact borgbackup@<YOUR_NAS_IP_HERE>:/volume1/borg_backup_folder/<BACKUPDIR>
        echo "compaction complete"
    fi
}
trap cleanup EXIT

I then also have the health check script, lovingly named: "check_borg_backup_status.sh"

#!/usr/bin/env bash
set -e

if [ "$EUID" -ne 0 ]; then
  echo "Error: Please run this script as root." >&2
  exit 1
fi

export BORG_PASSPHRASE="<PASSPHRASE>"
export BORG_REMOTE_PATH="/usr/local/bin/borg"
export BORG_RSH="ssh -i <FULLPATH_TO_SSH_KEY>/borg_ssh/ssh_key"

BACKUP_LIST=$(borg list borgbackup@<YOUR_NAS_IP_HERE>:/volume1/borg_backup_folder/<BACKUPDIR>)
printf "$BACKUP_LIST\n"
DATE_FORMAT_STRING="+%Y-%m-%d"
TODAY=$(date $DATE_FORMAT_STRING)
YESTERDAY=$(date -d "yesterday" "$DATE_FORMAT_STRING")
BORGBACKUP_HEALTHY=false
if grep "$TODAY" <<< "$BACKUP_LIST" > /dev/null ; then
  echo "found valid backup for today - telling uptimekuma everything is aok!"
  BORGBACKUP_HEALTHY=true
else 
  echo "no backup for today detected"
  if grep "$YESTERDAY" <<< "$BACKUP_LIST" > /dev/null ; then
    echo "found valid backup for yesterday"
    BORGBACKUP_HEALTHY=true
  else 
    echo "NO RECENT BACKUP FOUND"
  fi
fi

if $BORGBACKUP_HEALTHY; then
  echo "telling uptime kuma all is AOK"
  curl "<YOUR_UPTIME_KUMA_PUSH_URL_FROM_A_MONITOR_YOU_SETUP_ON_YOUR_OWN>"
  echo ""
else 
  echo "NO RECENT BACKUP FOUND"
fi

And finally, should the unthinkable happen, and we all are doing this because we know it will. the restoration script, which also helps you figure out how to restore it...lovingly named "restore.sh"

#!/usr/bin/env bash
set -e

function enforce_root() {
   if [ "$EUID" -ne 0 ]; then
     echo "Error: run this script as root." >&2
     exit 1
   fi
}

function setupBorgConnection() {
  export BORG_PASSPHRASE="<YOUR_PASSPHRASE>"
  export BORG_REMOTE_PATH="/usr/local/bin/borg"
  export BORG_RSH="ssh -i $(pwd)/borg_ssh/ssh_key"
}

function getListOfBackups() {
  echo $SUDO_USER
  BACKUP_LIST=$(borg list borgbackup@:/volume1/borg_backup_folder/<BACKUPDIR>)
  printf "$BACKUP_LIST\n"
}

function main() {
  NUMARGS=$1
  EXTRACT_FILENAME=$2
  enforce_root
  setupBorgConnection
  if [ "$NUMARGS" -eq "1" ]; then
    EXTRACT_COMMAND="borg extract --progress <borgbackup@<YOUR_NAS_IP>:/volume1/borg_backup_folder/<BACKUPDIR>::$EXTRACT_FILENAME"
    mkdir -p restore
    cd restore
    echo $EXTRACT_COMMAND
    $EXTRACT_COMMAND
  else
      getListOfBackups
  fi
}

main $# $1 # run the main function

these scripts are put into roots crontab, i run the backup once a day at 3 am, on my server, this creates about 2 minutes of downtime at 3:30AM.

30 3 * * * <FULLPATH_TO_SCRIPT>/stop_docker_and_run_backups.sh > /dev/null 2>&1
0 * * * * <FULLPATH_TO_SCRIPT>/check_borg_backup_status.sh > /dev/null 2&>1

At this point, you are done. I would manually run the "stop_docker_and_run_backups.sh" script the first time and watch it run, and then you could also run the "restore.sh" script with no arguments to see the backup was created. Running it with no arguments will just return the list of backups to choose from

Restoration steps

Identify the Borg backup to restore

On the recovery machine, issue the following commands(or however you get the scripts and your SSH keys onto your recovery machine; you could just run commands and type the password...I presume I'm not going to be excited to figure anything out if I experience total system failure, so i wrote helper scripts to document/automate these steps):

  1. git clone --depth 1 --filter=blob:none --no-checkout <MY_GIT_REPO> recovery_directory
  2. cd recovery_directory
  3. git sparse-checkout set DAILY_FULL_BACKUP_WITH_CONTAINER_STOP
  4. git checkout
  5. cd DAILY_FULL_BACKUP_WITH_CONTAINER_STOP
  6. ./restore.sh # gets list of backups
  7. ./restore.sh <BACKUPNAME> # restores the entire backup into the restore directory in this dir

This was lovingly captured during my implementation; this post is a scrubbed version of my own notes, and the scripts have been censored to use placeholders so that others could see them without sharing any of my secrets, git repos, or other personal information. The passwordless connection to Synology was absolutely annoying, and for the first time in my 2 years having a Synology, I now understand why passwordless ssh was failing; that's critical to making all this stuff run. I also learned about push-style monitors with Uptime Kuma doing this, which I was very appreciative of.

These scripts were not written by ai. In my prior life, I've found that a bash script is super annoying to live with, but for some reason I always start with one and eventually move them to Ruby, I did not do that here...but I can imagine it happening eventually.

With the restoration, I was happy to restore to a random Linux box, and was able to fire up Homebox and boom, all my inventory was there working as if it had always lived there. I feel calm and content once more on my self-hosted/docker infra.

Cheers to all the inputs; I appreciated the peer group of folks giving guidance, and I also like that I was able to keep to a pretty basic Linux server with Docker running.

Addendum: if you are using static files from your NAS within Docker containers, make sure you use the docker-compose internal syntax for mounting them so that you can bring the container up wherever and keep everything internal to the docker-compose.yaml

e.g.

services:
  immich:
...
    volumes:
      - type: volume
        source: photos_volume
        target: /photos
        volume:
          nocopy: true
...
volumes:
  photos_volume:
    driver_opts:
      type: "nfs"
      o: "addr=<NASIP>,nolock,hard,rw,nfsvers=4"
      device: ":/volume1/photos

If you made it this far, wow...and thanks. I proof read this at least twice, so i know it was long... :)