Nginx Docker容器被Systemd停止且无法自动重启的原因排查及预防方案咨询
Hey folks, let's dig into why your Nginx Docker container got stopped by systemd and wouldn't restart automatically, plus how to fix this and prevent future issues.
问题原因分析
Looking at your syslog and Nginx logs, we can break down what happened into two key parts:
1. Why was the Nginx container stopped?
The log clearly shows systemd initiated the stop:
Feb 28 06:44:43 elephantus systemd[1]: Stopping nginx container...
This happened right after a series of systemd service restarts (including systemd itself reexecuting, networkd, logind, etc.). The trigger here is almost certainly the daily apt upgrade/clean activity that kicked off at 06:44:08. System updates often trigger service reloads or restarts, and if your Nginx container is managed via a custom systemd unit file, systemd would have stopped it as part of this routine process.
2. Why didn't it restart automatically?
The critical error from your Nginx logs tells the full story:
2024/02/28 07:56:35 [emerg] 1#1: host not found in upstream "grafana" in /etc/nginx/conf.d/grafana.conf:15
When systemd tried to restart the Nginx container, it couldn't resolve the grafana upstream host. This means the Grafana container wasn't running (or wasn't reachable via Docker's internal DNS) at that moment. Nginx treats upstream resolution failures during startup as an emergency error, so it exits immediately. If your systemd unit or Docker restart policy wasn't configured to handle this failure and retry, the container stayed stopped until you manually intervened.
预防与修复方案
Here are actionable steps to fix this and stop it from recurring:
1. Let Docker handle container restarts (recommended)
If you're using Docker, leverage its built-in restart policies instead of relying solely on systemd. This ensures Docker will automatically restart containers if they exit, and can handle dependencies more reliably. Run these commands to update the restart policy for both Nginx and Grafana:
docker update --restart=always nginx docker update --restart=always grafana
--restart=always will restart the container no matter why it exited. If you prefer to let it stay stopped if you manually stop it, use --restart=unless-stopped instead.
2. Fix your systemd unit (if you need to keep using systemd)
If you're managing the container via a systemd unit file, tweak it to handle dependencies and retries:
- Add dependency rules to ensure Nginx starts only after Docker and Grafana are ready. In your nginx container's
.servicefile:After=docker.service grafana-container.service Requires=docker.service - Add restart logic to retry on failure:
Restart=on-failure RestartSec=5
This tells systemd to wait 5 seconds and retry starting Nginx if it fails, giving Grafana time to spin up.
3. Make Nginx more resilient to upstream failures
Modify your Nginx config to handle cases where Grafana isn't available during startup:
- Add a Docker DNS resolver to your upstream block, so Nginx can dynamically resolve the Grafana container even if it restarts later:
upstream grafana { server grafana:3000; resolver 127.0.0.11 valid=30s; # Docker's internal DNS fail_timeout 10s; } - Use
proxy_next_upstreamto handle temporary failures in your location block:location /grafana/ { proxy_pass http://grafana/; proxy_next_upstream error timeout invalid_header http_500 http_502 http_503 http_504; }
This way, Nginx will start successfully even if Grafana is down, and retry connecting automatically later.
4. Mitigate system update impacts
- Adjust the timing of daily apt updates to run during off-peak hours, so service restarts don't affect traffic. You can edit the apt timer file at
/etc/systemd/system/timers.target.wants/apt-daily-upgrade.timer. - Configure apt to avoid automatic service restarts during updates (use with caution—this might leave some services in a pending restart state). Add
APT::Periodic::Unattended-Upgrade::AutoFixInterruptedDpkg "false";to/etc/apt/apt.conf.d/50unattended-upgrades.
备注:内容来源于stack exchange,提问作者Alakananda S

