Systemd定时器触发ZFS快照S3备份脚本异常问题求助
Let's break down your tricky problem: your backup script works flawlessly when run manually as root or started via systemctl start, but when triggered by a Systemd Timer, the uploaded file ends up empty after decryption—even though the upload duration matches what you'd expect for 900GB of data. That's definitely counterintuitive, so let's walk through the most likely causes and fixes.
1. Environment Variable Mismatches
Systemd services (especially those triggered by timers) run with a stripped-down default environment—far fewer variables than your interactive root shell. This could break variables like BACKUP_NAME, BACKUP_PASS, or even paths to tools like zfs, lz4c, or your $s3_command.
Fix Steps:
- Log environment variables for comparison: Add this line at the start of your
/root/s3backup.shscript:
Run the script manually, then trigger it via the timer, and compare the two log files. Look for differences inenv > /tmp/backup_env_$(date +%Y%m%d_%H%M%S).logPATH,HOME,BACKUP_NAME, orBACKUP_PASS. - Explicitly set required variables in your Systemd service unit. Either add individual
Environmentlines or load from a secure file:
Make sure# In your service unit EnvironmentFile=/root/backup.env # File containing BACKUP_NAME=..., BACKUP_PASS=.../root/backup.envhas strict permissions (chmod 600 /root/backup.env) to keep credentials secure.
2. Missing or Invalid ZFS Snapshot
If the snapshot ztank/data@$BACKUP_NAME doesn't exist when the timer runs, zfs send will output nothing—but it might hang instead of exiting immediately, leading to that confusing long upload duration.
Fix Steps:
- Add a snapshot existence check at the start of your script:
SNAPSHOT="ztank/data@$BACKUP_NAME" if ! zfs list -H "$SNAPSHOT" > /dev/null 2>&1; then echo "ERROR: Snapshot $SNAPSHOT does not exist!" >> /tmp/backup_errors.log exit 1 fi - Verify snapshot creation logic: If your script creates the snapshot before sending it, add logging to that step to confirm it works in the timer context:
zfs snapshot "$SNAPSHOT" >> /tmp/backup_log.log 2>&1
3. Silent Failures in the Pipeline
By default, Bash only checks the exit code of the last command in a pipeline. If zfs send or openssl fails silently, the rest of the pipeline will process empty input, and s3 cp will upload an empty file without triggering an error.
Fix Steps:
- Enable strict error checking in your script. Add this at the top:
This makes the script exit immediately if any command in the pipeline fails, and propagates the error code.set -eo pipefail - Log each pipeline step to isolate where data is getting lost:
After the timer runs, check these logs. Ifzfs send -c "$SNAPSHOT" > /tmp/zfs_send.log 2>&1 \ | openssl enc -aes-256-cbc -iter 1000 -salt -pass pass:"$BACKUP_PASS" > /tmp/openssl.log 2>&1 \ | lz4c > /tmp/lz4c.log 2>&1 \ | "$s3_command" s3 cp --expected-size 188743680 - s3://instamongo.backup/"$BACKUP_NAME".lz/tmp/zfs_send.logis empty,zfs sendis the problem. Ifzfs_send.loghas data butopenssl.logshows errors, encryption is failing.
4. Systemd Service Execution Context
Even with User=root, the timer-triggered service might have a different working directory or resource constraints than when you run it manually.
Fix Steps:
- Explicitly set working directory and PATH in your service unit:
# In your service unit WorkingDirectory=/root Environment="PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin" - Check Systemd logs for hidden errors: Run
journalctl -u backup-mongo.service -fto watch logs in real time when the timer triggers. Look for warnings or errors that might not be printed to your script's logs.
5. Pipeline Blocking from Resource Locks
That long upload duration suggests something is hanging in the pipeline. It's possible zfs send is stuck waiting for a locked dataset when triggered by the timer, but never outputs data. Check dataset locks with zfs get locks ztank/data after a failed timer run to confirm this.
内容的提问来源于stack exchange,提问作者gdanov

