Ubuntu 21.04环境下nohup运行的Python无限循环脚本运行数天后意外停止问题排查
Hey there! Let's break down your problem and figure out why your Python script keeps stopping after ~3 days, and how to fix it.
First off: Ubuntu 21.04 doesn’t have an inherent "kill long-running background processes" feature—so the issue is almost certainly coming from somewhere else, not the OS itself. Here are the most likely culprits and solutions:
1. GCP VM Maintenance Events
Google Cloud regularly performs maintenance on its infrastructure (like kernel updates or hardware checks). If your VM is set to the default "Migrating" or "Terminating" policy for maintenance, it might restart, which would kill your nohup-managed process.
- Check: Head to your GCP Console → VM Instances → select your VM → look at the "Operations" tab to see if there’s a restart or maintenance event around the time your script died.
- Fix: Set your VM to auto-restart (under VM settings → "Availability policies") and use a process manager like systemd (more on that below) to automatically restart the script after a reboot.
2. Unhandled Signals or Silent Crashes
Even with nohup, your script might be receiving termination signals (like SIGTERM, SIGHUP from unexpected system events) or crashing silently without logging the issue.
Capture Signals: Add signal handling to your script to log any incoming termination signals. This will help you pinpoint if a signal is killing the process:
import signal import logging import sys # Set up logging to track events logging.basicConfig( filename='script_events.log', level=logging.INFO, format='%(asctime)s - %(message)s' ) def handle_signal(signum, frame): signal_name = signal.Signals(signum).name logging.info(f"Script received signal: {signal_name} ({signum})") # Add cleanup logic here if needed, then exit gracefully sys.exit(0) # Register handlers for common termination signals signal.signal(signal.SIGTERM, handle_signal) signal.signal(signal.SIGHUP, handle_signal) signal.signal(signal.SIGINT, handle_signal)Capture All Output: Your current
nohupcommand only redirects stdout—stderr (where Python exceptions are logged) might be going nowhere. Update your launch command to capture both:nohup python3 -u main.py > output.log 2>&1 &This merges stderr into the same log file, so you’ll see any uncaught exceptions that might be crashing the script.
3. Resource Leaks (Not Just OOM)
While you don’t see "Killed" in logs, your script might be hitting other resource limits over time:
- File Descriptors: If your task opens files/databases without properly closing them, you could hit the system’s file descriptor limit. Use Python’s
withstatement for file operations to auto-close resources. - Memory Fragmentation: Even if you’re not using enough memory to trigger OOM, long-running Python scripts can accumulate fragmented memory that causes unexpected crashes. Adding periodic garbage collection (
import gc; gc.collect()) might help, but fixing resource leaks is better.
4. Ditch nohup for systemd
nohup is a quick-and-dirty way to run background processes, but it’s not designed for long-term reliability. systemd is the standard process manager on Ubuntu, and it will automatically restart your script if it crashes, log all output, and handle system reboots.
Here’s how to set it up:
- Create a systemd service file (e.g.,
/etc/systemd/system/periodic-task.service):[Unit] Description=My Periodic Python Task After=network.target [Service] User=your-username # Replace with your actual username WorkingDirectory=/path/to/your/script/folder ExecStart=/usr/bin/python3 /path/to/your/script/main.py Restart=always # Auto-restart on crash or reboot RestartSec=60 # Wait 60 seconds before restarting StandardOutput=append:/var/log/periodic-task.log StandardError=append:/var/log/periodic-task.log [Install] WantedBy=multi-user.target - Enable and start the service:
sudo systemctl daemon-reload sudo systemctl enable periodic-task.service sudo systemctl start periodic-task.service - Check status or logs anytime:
sudo systemctl status periodic-task.service sudo tail -f /var/log/periodic-task.log
Final Takeaway
Ubuntu 21.04 isn’t the culprit here. The most likely causes are GCP maintenance, unlogged crashes, or unreliable process management with nohup. Using systemd will make your script far more resilient, and adding signal handling + full output logging will help you diagnose any remaining issues quickly.
内容的提问来源于stack exchange,提问作者Jeffrey Chen

