You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Kubernetes滚动更新或副本集缩容时为何出现服务中断?

Hey there, let's break down what's happening here and how to fix it!

What "Zero Downtime" Actually Means in Kubernetes

First, let's clear up the confusion: Kubernetes' "zero downtime" guarantee doesn't mean every single request will succeed, even under extreme load like your 0.01-second interval curls. Instead, it's designed to prevent extended periods of total service unavailability and ensure the vast majority of requests are handled normally. Your test is an edge case with ultra-frequent, nearly continuous requests, which exposes some subtle behaviors that aren't usually noticeable in real-world production workloads.

Why You're Seeing Connection Resets

Let's walk through the specific reasons your requests are failing during updates and scaling:

1. Aggressive Rolling Update Strategy

Your deployment uses maxSurge: 0 and maxUnavailable: 9. This means Kubernetes will delete 9 out of 10 old pods first before creating any new ones. During this window, you only have 1 pod handling all traffic—way too few for your ultra-high request rate. Worse, when old pods are terminated, any in-flight connections to them get immediately reset because there's no grace period for existing requests to finish.

Even the default strategy (25% max surge/unavailable) might show occasional failures in your test, but your current configuration amplifies the problem drastically.

2. Pod Termination Doesn't Handle Existing Connections Gracefully

When Kubernetes deletes a pod, it does two things:

  • Removes the pod from the Service's Endpoint list (so new requests don't go to it)
  • Sends a TERM signal to the container to shut down

By default, Tomcat doesn't handle the TERM signal gracefully—it will immediately close its listening port and terminate all active connections, causing the "Connection reset by peer" error you see. Plus, there's a small delay between removing the pod from Endpoints and kube-proxy updating its routing rules (especially with iptables mode in minikube), so some requests might still route to the terminating pod.

3. Scaling Down Triggers the Same Termination Behavior

When you scale from 10 to 2 replicas, Kubernetes deletes 8 pods at once. Again, those terminating pods drop active connections, and kube-proxy's routing updates can't keep up with your 100 requests per second.

Fixes to Minimize or Eliminate These Interrupts

Here's what you can do to fix this:

1. Adjust Your Rolling Update Strategy

Use a more conservative strategy that prioritizes keeping pods available:

strategy:
  type: RollingUpdate
  rollingUpdate:
    maxSurge: 2  # Start 2 new pods before deleting old ones
    maxUnavailable: 0  # Never let the number of available pods drop below the desired count

This way, you'll always have at least 10 pods handling traffic during updates. Kubernetes will spin up 2 new pods, wait for them to pass readiness checks, then delete 2 old pods—repeat until all are updated.

2. Configure Graceful Termination for Tomcat

  • Add a terminationGracePeriodSeconds to your pod template (default is 30s, but you can increase it if needed):
    spec:
      terminationGracePeriodSeconds: 60
      containers:
      - name: tc-part
        image: tomcat:8-jdk11
        # ... rest of your config
    
  • Modify Tomcat to handle the TERM signal gracefully:
    • Create a custom entrypoint script that listens for TERM and runs $CATALINA_HOME/bin/shutdown.sh instead of letting the container exit immediately. The shutdown script tells Tomcat to stop accepting new connections and wait for existing ones to finish before exiting.
    • Alternatively, add a lifecycle hook to your container:
      lifecycle:
        preStop:
          exec:
            command: ["/bin/sh", "-c", "$CATALINA_HOME/bin/shutdown.sh"]
      

3. Optimize Kube-Proxy Routing

Minikube uses the iptables mode for kube-proxy by default, which has a small delay when updating routing rules. For faster Endpoint updates, you can switch to IPVS mode (available in Kubernetes 1.11+):

minikube start --extra-config=kube-proxy.mode=ipvs

IPVS handles connection routing and Endpoint updates more efficiently, reducing the chance of requests hitting terminating pods.

4. Add Client-Side Retries

Even with all server-side optimizations, ultra-high request rates might still see occasional failures. Make your client retry idempotent requests (like your GET to /) when it hits connection errors. For example, modify your curl loop to retry on failure:

for ((;;)); do
  curl -sS -D - http://tc-webapp-service:1234 -o /dev/null | grep HTTP || echo "Request failed, retrying..."
  date +"%Y-%m-%d %H:%M:%S"
  echo 
  sleep 0.01 
done;

5. Refine Your Readiness Probe

Ensure your readiness probe truly reflects when the pod is ready to handle traffic. Instead of just hitting /, use a path that checks if Tomcat and your web app are fully initialized. For example:

readinessProbe:
  httpGet:
    scheme: HTTP
    port: 8080
    path: /manager/html  # Or a custom health check endpoint for your app
  initialDelaySeconds: 10
  periodSeconds: 5

This prevents new pods from being added to the Service until they're actually ready to process requests.


内容的提问来源于stack exchange,提问作者absuu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:25:35