GKE部署的Django应用滚动更新时MySQL连接失败问题及解决方案咨询
I’ve run into this exact issue before—during GKE rolling updates, the Cloud SQL Proxy container shuts down first, leaving the Django app still running and unable to connect to the database. The fix boils down to controlling the shutdown order and timing of your containers to make sure Django stops before the proxy does. Here are two proven solutions:
Option 1: Add a Termination Timeout to Cloud SQL Proxy (Quickest Fix)
The Cloud SQL Proxy has a built-in -term-timeout flag that tells it to wait a specified amount of time after receiving a TERM signal before shutting down. This gives your Django app enough time to finish handling requests and exit gracefully.
Update your myapp.yaml Cloud SQL Proxy container command to include this flag:
- image: gcr.io/cloudsql-docker/gce-proxy:1.16 name: cloudsql-proxy command: ["/cloud_sql_proxy", "--dir=/cloudsql", "-instances=myproject:europe-north1:myapp=tcp:3306", "-credential_file=/secrets/cloudsql/credentials.json", "-term-timeout=30s"] # Add this line to wait 30 seconds before exiting
With this, the proxy will keep running for 30 seconds after getting the shutdown signal, giving Django plenty of time to wrap up and exit without hitting DB connection errors.
Option 2: Add a PreStop Hook to Django (More Elegant)
If you want Django to actively stop accepting new requests and finish existing connections before shutting down, add a preStop hook to your app container, paired with an extended terminationGracePeriodSeconds:
Modify the Django container section in your Deployment:
- name: myapp-app image: gcr.io/myproject/myapp imagePullPolicy: IfNotPresent lifecycle: preStop: exec: command: ["sleep", "10"] # Wait 10 seconds to let the app wind down terminationGracePeriodSeconds: 40 # Total shutdown window (needs to be longer than preStop)
How This Works
Kubernetes follows this sequence when terminating a Pod:
- Removes the Pod from the Service’s endpoint list (no new requests will reach it)
- Sends a TERM signal to all containers
- Runs any
preStophooks defined for containers - Waits
terminationGracePeriodSeconds; sends a KILL signal if containers haven’t exited by then
The sleep 10 in the preStop hook gives Django time to:
- Finish processing any in-flight requests
- Close DB connections cleanly
- Exit normally before the proxy shuts down
If you’re using a process manager like Gunicorn, you can replace the sleep with a command to trigger graceful shutdown directly:
command: ["pkill", "-SIGTERM", "gunicorn"] # Send SIGTERM to Gunicorn to initiate graceful shutdown
Bonus: Add Health Checks for Smoother Updates
To make your rolling updates even more reliable, add readiness and liveness probes to your Django container. Kubernetes will wait for the new Pod to be fully ready before terminating old ones:
- name: myapp-app image: gcr.io/myproject/myapp imagePullPolicy: IfNotPresent livenessProbe: httpGet: path: /api/health/ port: 8080 initialDelaySeconds: 30 periodSeconds: 10 readinessProbe: httpGet: path: /api/health/ port: 8080 initialDelaySeconds: 5 periodSeconds: 5 # Include your preStop hook and terminationGracePeriodSeconds here
This ensures that only fully functional Pods receive traffic, minimizing downtime during updates.
内容的提问来源于stack exchange,提问作者Egor Wexler

