Gunicorn内存占用与线程持续增长问题排查求助
Hey there, let's break down why you're seeing those lingering Gunicorn processes and skyrocketing memory usage in your Kubernetes Django setup—this is a pretty common pitfall, so we'll walk through the fixes step by step.
1. The --reload flag is dangerous in production
First off, that --reload flag you're using is meant only for local development. Here's why it's causing chaos in production:
- It spins up an extra monitoring process that watches for file changes to restart workers. In Kubernetes, even minor filesystem blips (like ConfigMap/Secret updates, temporary file writes) can trigger constant worker restarts.
- The reload mechanism isn't built for production-grade process management—old worker processes often don't get properly cleaned up when new ones start, leaving zombie processes hanging around and hogging memory. These zombies are dead but not reaped by the parent Gunicorn process, so their memory stays allocated.
2. Potential process reaping failures
Even without --reload, if the Gunicorn master process doesn't handle worker exit signals correctly (like when Kubernetes sends termination signals or hits memory limits), you can get leftover zombies. But combined with --reload, this is almost certainly the main culprit.
1. Ditch the --reload flag immediately
This is the quickest win. Your production Gunicorn command shouldn't include --reload—replace it with:
gunicorn app.wsgi -w 3 -b 0.0.0.0:8000 --env DJANGO_SETTINGS_MODULE=app.settings.prod
This runs Gunicorn in pure production mode, no extra monitoring process, and uses its stable, battle-tested worker management logic.
2. Add worker recycling rules to prevent memory leaks
Even with a clean production setup, Django apps can slowly leak memory over time. Configure Gunicorn to automatically recycle workers to keep memory in check:
--max-requests N: Restart each worker after it handles N requests (prevents memory buildup from long-running workers)--max-requests-jitter N: Add randomness to the restart threshold so all workers don't reboot at the same time (avoids service downtime)--timeout 30: Kill workers that hang on requests longer than 30 seconds
Updated command example:
gunicorn app.wsgi -w 3 -b 0.0.0.0:8000 --env DJANGO_SETTINGS_MODULE=app.settings.prod --max-requests 1000 --max-requests-jitter 200 --timeout 30
3. Hardening your Kubernetes Deployment
Add resource limits and health probes to Kubernetes to catch and fix issues before they spiral:
- Set memory limits to force Pod restarts if memory gets too high
- Liveness/readiness probes make sure Kubernetes only sends traffic to healthy Pods, and restarts unresponsive ones
Here's a snippet for your Deployment YAML:
apiVersion: apps/v1 kind: Deployment metadata: name: django-app spec: replicas: 3 template: spec: containers: - name: django-app image: your-app-image:latest command: ["gunicorn"] args: [ "app.wsgi", "-w", "3", "-b", "0.0.0.0:8000", "--env", "DJANGO_SETTINGS_MODULE=app.settings.prod", "--max-requests", "1000", "--max-requests-jitter", "200", "--timeout", "30" ] resources: requests: memory: "256Mi" # Minimum memory allocated to the Pod limits: memory: "512Mi" # Kill and restart if memory exceeds this livenessProbe: httpGet: path: /healthz # You'll need to add a health check endpoint in Django port: 8000 initialDelaySeconds: 30 # Give the app time to start periodSeconds: 10 # Check every 10 seconds readinessProbe: httpGet: path: /healthz port: 8000 initialDelaySeconds: 5 periodSeconds: 5
Note: You'll need to add a simple /healthz view in your Django app that returns a 200 OK response when the app is healthy—this lets Kubernetes verify your service is working.
4. Check for Django app-level memory leaks
If you still see memory creeping up after fixing Gunicorn and Kubernetes configs, the issue might be in your Django code:
- Use tools like
memory_profilerto track which parts of your code are using the most memory - Look for global variables that accumulate data over time, unclosed database connections/file handles, or leaky third-party libraries
If your Pods are already using too much memory and causing issues, restart the Deployment to wipe all old processes:
kubectl rollout restart deployment/django-app
内容的提问来源于stack exchange,提问作者kvnm

