Kubernetes滚动更新或副本集缩容时为何出现服务中断?
Hey there, let's break down what's happening here and how to fix it!
What "Zero Downtime" Actually Means in Kubernetes
First, let's clear up the confusion: Kubernetes' "zero downtime" guarantee doesn't mean every single request will succeed, even under extreme load like your 0.01-second interval curls. Instead, it's designed to prevent extended periods of total service unavailability and ensure the vast majority of requests are handled normally. Your test is an edge case with ultra-frequent, nearly continuous requests, which exposes some subtle behaviors that aren't usually noticeable in real-world production workloads.
Why You're Seeing Connection Resets
Let's walk through the specific reasons your requests are failing during updates and scaling:
1. Aggressive Rolling Update Strategy
Your deployment uses maxSurge: 0 and maxUnavailable: 9. This means Kubernetes will delete 9 out of 10 old pods first before creating any new ones. During this window, you only have 1 pod handling all traffic—way too few for your ultra-high request rate. Worse, when old pods are terminated, any in-flight connections to them get immediately reset because there's no grace period for existing requests to finish.
Even the default strategy (25% max surge/unavailable) might show occasional failures in your test, but your current configuration amplifies the problem drastically.
2. Pod Termination Doesn't Handle Existing Connections Gracefully
When Kubernetes deletes a pod, it does two things:
- Removes the pod from the Service's Endpoint list (so new requests don't go to it)
- Sends a
TERMsignal to the container to shut down
By default, Tomcat doesn't handle the TERM signal gracefully—it will immediately close its listening port and terminate all active connections, causing the "Connection reset by peer" error you see. Plus, there's a small delay between removing the pod from Endpoints and kube-proxy updating its routing rules (especially with iptables mode in minikube), so some requests might still route to the terminating pod.
3. Scaling Down Triggers the Same Termination Behavior
When you scale from 10 to 2 replicas, Kubernetes deletes 8 pods at once. Again, those terminating pods drop active connections, and kube-proxy's routing updates can't keep up with your 100 requests per second.
Fixes to Minimize or Eliminate These Interrupts
Here's what you can do to fix this:
1. Adjust Your Rolling Update Strategy
Use a more conservative strategy that prioritizes keeping pods available:
strategy: type: RollingUpdate rollingUpdate: maxSurge: 2 # Start 2 new pods before deleting old ones maxUnavailable: 0 # Never let the number of available pods drop below the desired count
This way, you'll always have at least 10 pods handling traffic during updates. Kubernetes will spin up 2 new pods, wait for them to pass readiness checks, then delete 2 old pods—repeat until all are updated.
2. Configure Graceful Termination for Tomcat
- Add a
terminationGracePeriodSecondsto your pod template (default is 30s, but you can increase it if needed):spec: terminationGracePeriodSeconds: 60 containers: - name: tc-part image: tomcat:8-jdk11 # ... rest of your config - Modify Tomcat to handle the
TERMsignal gracefully:- Create a custom entrypoint script that listens for
TERMand runs$CATALINA_HOME/bin/shutdown.shinstead of letting the container exit immediately. The shutdown script tells Tomcat to stop accepting new connections and wait for existing ones to finish before exiting. - Alternatively, add a lifecycle hook to your container:
lifecycle: preStop: exec: command: ["/bin/sh", "-c", "$CATALINA_HOME/bin/shutdown.sh"]
- Create a custom entrypoint script that listens for
3. Optimize Kube-Proxy Routing
Minikube uses the iptables mode for kube-proxy by default, which has a small delay when updating routing rules. For faster Endpoint updates, you can switch to IPVS mode (available in Kubernetes 1.11+):
minikube start --extra-config=kube-proxy.mode=ipvs
IPVS handles connection routing and Endpoint updates more efficiently, reducing the chance of requests hitting terminating pods.
4. Add Client-Side Retries
Even with all server-side optimizations, ultra-high request rates might still see occasional failures. Make your client retry idempotent requests (like your GET to /) when it hits connection errors. For example, modify your curl loop to retry on failure:
for ((;;)); do curl -sS -D - http://tc-webapp-service:1234 -o /dev/null | grep HTTP || echo "Request failed, retrying..." date +"%Y-%m-%d %H:%M:%S" echo sleep 0.01 done;
5. Refine Your Readiness Probe
Ensure your readiness probe truly reflects when the pod is ready to handle traffic. Instead of just hitting /, use a path that checks if Tomcat and your web app are fully initialized. For example:
readinessProbe: httpGet: scheme: HTTP port: 8080 path: /manager/html # Or a custom health check endpoint for your app initialDelaySeconds: 10 periodSeconds: 5
This prevents new pods from being added to the Service until they're actually ready to process requests.
内容的提问来源于stack exchange,提问作者absuu

