You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Gunicorn内存占用与线程持续增长问题排查求助

Hey there, let's break down why you're seeing those lingering Gunicorn processes and skyrocketing memory usage in your Kubernetes Django setup—this is a pretty common pitfall, so we'll walk through the fixes step by step.

Root Cause Analysis

1. The --reload flag is dangerous in production

First off, that --reload flag you're using is meant only for local development. Here's why it's causing chaos in production:

  • It spins up an extra monitoring process that watches for file changes to restart workers. In Kubernetes, even minor filesystem blips (like ConfigMap/Secret updates, temporary file writes) can trigger constant worker restarts.
  • The reload mechanism isn't built for production-grade process management—old worker processes often don't get properly cleaned up when new ones start, leaving zombie processes hanging around and hogging memory. These zombies are dead but not reaped by the parent Gunicorn process, so their memory stays allocated.

2. Potential process reaping failures

Even without --reload, if the Gunicorn master process doesn't handle worker exit signals correctly (like when Kubernetes sends termination signals or hits memory limits), you can get leftover zombies. But combined with --reload, this is almost certainly the main culprit.

Fixes to Implement

1. Ditch the --reload flag immediately

This is the quickest win. Your production Gunicorn command shouldn't include --reload—replace it with:

gunicorn app.wsgi -w 3 -b 0.0.0.0:8000 --env DJANGO_SETTINGS_MODULE=app.settings.prod

This runs Gunicorn in pure production mode, no extra monitoring process, and uses its stable, battle-tested worker management logic.

2. Add worker recycling rules to prevent memory leaks

Even with a clean production setup, Django apps can slowly leak memory over time. Configure Gunicorn to automatically recycle workers to keep memory in check:

  • --max-requests N: Restart each worker after it handles N requests (prevents memory buildup from long-running workers)
  • --max-requests-jitter N: Add randomness to the restart threshold so all workers don't reboot at the same time (avoids service downtime)
  • --timeout 30: Kill workers that hang on requests longer than 30 seconds

Updated command example:

gunicorn app.wsgi -w 3 -b 0.0.0.0:8000 --env DJANGO_SETTINGS_MODULE=app.settings.prod --max-requests 1000 --max-requests-jitter 200 --timeout 30

3. Hardening your Kubernetes Deployment

Add resource limits and health probes to Kubernetes to catch and fix issues before they spiral:

  • Set memory limits to force Pod restarts if memory gets too high
  • Liveness/readiness probes make sure Kubernetes only sends traffic to healthy Pods, and restarts unresponsive ones

Here's a snippet for your Deployment YAML:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: django-app
spec:
  replicas: 3
  template:
    spec:
      containers:
      - name: django-app
        image: your-app-image:latest
        command: ["gunicorn"]
        args: [
          "app.wsgi",
          "-w", "3",
          "-b", "0.0.0.0:8000",
          "--env", "DJANGO_SETTINGS_MODULE=app.settings.prod",
          "--max-requests", "1000",
          "--max-requests-jitter", "200",
          "--timeout", "30"
        ]
        resources:
          requests:
            memory: "256Mi"  # Minimum memory allocated to the Pod
          limits:
            memory: "512Mi"  # Kill and restart if memory exceeds this
        livenessProbe:
          httpGet:
            path: /healthz  # You'll need to add a health check endpoint in Django
            port: 8000
          initialDelaySeconds: 30  # Give the app time to start
          periodSeconds: 10  # Check every 10 seconds
        readinessProbe:
          httpGet:
            path: /healthz
            port: 8000
          initialDelaySeconds: 5
          periodSeconds: 5

Note: You'll need to add a simple /healthz view in your Django app that returns a 200 OK response when the app is healthy—this lets Kubernetes verify your service is working.

4. Check for Django app-level memory leaks

If you still see memory creeping up after fixing Gunicorn and Kubernetes configs, the issue might be in your Django code:

  • Use tools like memory_profiler to track which parts of your code are using the most memory
  • Look for global variables that accumulate data over time, unclosed database connections/file handles, or leaky third-party libraries
Quick Emergency Fix

If your Pods are already using too much memory and causing issues, restart the Deployment to wipe all old processes:

kubectl rollout restart deployment/django-app

内容的提问来源于stack exchange,提问作者kvnm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 18:27:28