超54万图片(占500万Inode)每日异地高效备份方案咨询
Great question—dealing with millions of inodes and hundreds of thousands of small image files is a classic backup pain point. Rsync is a reliable tool, but its default behavior can drag with your scale because it has to stat every single file to check for changes. Let’s dive into optimizing rsync first, then cover faster alternatives that might be a better fit for your setup.
Optimizing Rsync for Large Inode Counts
If you want to stick with rsync, these tweaks will cut down your backup time significantly:
- Archive mode + selective compression: Use
-a(archive) to preserve permissions, timestamps, and directory structure. Add-z(compression) only if your network is slow—skip it for fast networks or uncompressible files like JPEGs, since it wastes CPU cycles.rsync -az /path/to/local/images user@remote:/path/to/backup/ - Skip incremental recursion: For rsync 3.0.0+,
--no-inc-recursiveforces rsync to scan the entire directory tree first before transferring files. This avoids repeated stat calls during traversal, which is a huge time-saver with 5 million inodes. - Incremental backups with hard links: The
--link-destflag is a game-changer for daily backups. Instead of re-transferring every file, rsync creates hard links to unchanged files from your last backup. This saves both transfer time and remote storage space.
Here’s a script example for daily timestamped backups:# Set today's date for the backup directory BACKUP_DATE=$(date +%Y%m%d) # Sync new/changed files, linking to the latest backup for unchanged ones rsync -az --link-dest=/path/to/backup/latest /path/to/local/images user@remote:/path/to/backup/$BACKUP_DATE # Update the 'latest' symlink to point to today's backup ssh user@remote "ln -sf /path/to/backup/$BACKUP_DATE /path/to/backup/latest" - Trim unnecessary metadata: If you don’t need to preserve permissions or ownership on the remote server, add
--no-permsor--no-ownerto cut down on metadata transfer overhead.
Faster Alternatives to Rsync for This Scenario
If even optimized rsync is too slow, these approaches work better with large numbers of small files:
Filesystem-Level Snapshots (Btrfs/ZFS)
If your local or remote filesystem uses Btrfs or ZFS, snapshot-based backups are the fastest option. They operate at the block level, so you don’t have to stat millions of inodes—you just send the changes since your last snapshot.
- Btrfs example:
# Create a read-only snapshot of your images directory btrfs subvolume snapshot -r /path/to/images /path/to/temp_snapshot # Send the snapshot to the remote server (use -i for incremental transfers after the first backup) btrfs send /path/to/temp_snapshot | ssh user@remote "btrfs receive /path/to/remote/backup" # Clean up the local snapshot once the transfer is done btrfs subvolume delete /path/to/temp_snapshot - ZFS example:
# Create a snapshot named after today's date zfs create pool/images@daily_$(date +%Y%m%d) # Send the incremental changes from the last snapshot to the remote zfs send -i pool/images@last_backup pool/images@daily_$(date +%Y%m%d) | ssh user@remote "zfs receive pool/remote_backup" # Rename today's snapshot to 'last_backup' for tomorrow's incremental transfer zfs rename pool/images@daily_$(date +%Y%m%d) pool/images@last_backup
This method is drastically faster than rsync for your scale because it avoids file-by-file traversal entirely.
Tar + Rsync for Bulk Transfers
If you can’t use filesystem snapshots, bundling rarely changed directories into tar archives reduces the number of files rsync has to process. Use incremental tar to only include changed files each day:
# Create an incremental tar archive (only adds files changed since the last snapshot) tar -cf /tmp/images_backup.tar --listed-incremental=/tmp/snapshot.snar /path/to/images # Sync the tar file to the remote server rsync -az /tmp/images_backup.tar user@remote:/path/to/backup/ # Extract the archive on the remote if needed ssh user@remote "tar -xf /path/to/backup/images_backup.tar -C /path/to/restore/"
The --listed-incremental flag keeps the archive size small by only including new/modified files.
Rclone for Cloud/Object Storage
If your remote is a cloud service (like S3, Google Drive), rclone is optimized for bulk transfers and handles large numbers of small files better than rsync. It supports parallel transfers to speed up uploads:
# Sync local images to remote cloud storage with increased parallelism rclone sync /path/to/images remote:backup/images --transfers 32 --checkers 64
The --transfers flag controls how many files are uploaded at once, and --checkers increases the number of parallel file checks.
Additional Performance Boosters
- Network tweaks: Use a wired connection instead of Wi-Fi. If using SSH, enable compression in
sshd_config(addCompression yes) to reduce transfer size. - Archive static files: If some directories haven’t changed in months, archive them into a single tar file. This reduces the number of inodes rsync needs to scan daily.
- Parallelize rsync: Tools like
parsyncor custom scripts that sync directories in parallel can speed up transfers for deep tree structures.
内容的提问来源于stack exchange,提问作者raullarion

