Comment on Release v3.3.0 · immich-app/immich
avidamoeba@lemmy.ca 23 hours ago Immich has so many dirs and files that it will take hour/hours for any file-based program to walk that tree (on an HDD) in order to do a backup. I used to use rsync, syncthing, duplicity and had gotten it to about under an hour for 1TB library. The service has to be shutdown to preserve consistency of the backup. So that downtime was okay once per day. Still a heavy operation. Not great, not terrible.
Because of the above you might want to do one of the following:
- Use snapshotting filesystem like ZFS or Btrfs. That allows to avoid the downtime as you can backup a frozen-in-time snapshot while Immich is running.
- The above but with ZFS om both the Immich end and the backup end. That allows transferring just the delta between the new Immich snapshot and the last snapshot the backup has. This happens entirely within the filesystem so there are no file operations needed. At every point in time ZFS knows what blocks have changed between now and the previously backed up snapshot. When you ask it for that, it starts spitting data without spensing an hour figuring out what changed. That allows online backups every hour or even more frequently. The backup of my 1TB Immich instance (HDD) takes 5-10 seconds to initiate the network transfer, then it takes whatever time it takes to get to the backup machine, which is usually equivalent to transferring whatever new was stored in Immich.
So depending on how painful migrations are for you, I’d highly recommend the latter. Otherwise the former. The ZFS snapshot send strategy works for any service you may run.
Immich uses YYYY/MM folders on my setup, and I exclude any thumbnail directories from the backup too. So not many dirs/files to check for the backup software.
You don’t need to shut it down, pg_dump works fine while it’s like if you want to make sure it’s a good database backup. The files side of things are fine either way.
That makes sense. I copy the whole set of data dirs so I can trivially start it after restore or start it elsewhere without extra steps. Also because this strategy works with all services so I don’t have to consider how to backup/restore each one. Makes adding new services less work.
pgdump with live file copy while the service is in active use can result in files the db doesn’t know about, or files it thinks are there that were actuall deleted. Probably can be fixed after the fact.
More generally, as a someone who’s done software for a very long time, I’ve learned that the further away I go from the happy path of a software program, the less tested it is, the more bugs there are and the poorer the edge case handling is. So for backups I lean on the process kill/recovery edge case that they all must handle. Snapshot + backup from that snapshot looks like a process death to the service upon restore.
That’s fair, shutdown service + backup everything is certainly more guaranteed to work properly.
With Immich I’m not too worried about the DB potentially being off since it can rescan the filesystem, but I also backup at 3am when nothing is happening on Immich because I’m asleep!
If you do it at 3AM anyway… may as well eliminate the need to even think about mismatches. 😁 I think mine also use to do it at 3AM before I switched to ZFS snapshots + send/recv.