Kubernetes环境下Tomcat探针配置失败问题咨询
Let's break down the best practices for configuring liveness and readiness probes for Tomcat-based apps on Kubernetes, plus fix the probe failures you're seeing.
First, let's diagnose why your current setup might be failing:
- You're using
/as the probe path, but if your app isn't deployed to Tomcat's ROOT context, or the root path returns a non-200 status (like 404 if you replaced the default Tomcat welcome page), the probe will fail. - Your
initialDelaySeconds(20s) might be too short—Tomcat plus your Java app's initialization could take longer than that, leading to premature failed checks.
Best Practices for Tomcat Probes
1. Use Your App's Custom Health Check Endpoint (Preferred)
If your Java web app has a built-in health check (like /actuator/health for Spring Boot apps), this is the best choice. It doesn't just check if Tomcat is running—it verifies your app's internal state (database connections, cache health, etc.).
2. Use Tomcat's Built-in Status Endpoint (If No Custom Check)
Tomcat includes a /manager/status endpoint that reports the server's runtime state. To use it, first configure Tomcat to allow access:
- Add a user with the right roles in
conf/tomcat-users.xml:
<user username="probe-user" password="secure-probe-pass" roles="manager-status"/>
- Then include authentication in your probe config (more on that below).
3. Tune Probe Parameters Wisely
- initialDelaySeconds: Set this based on your app's actual startup time. Check pod logs with
kubectl logs <pod-name>to see when your app is fully ready—start with 60s if you're unsure, then adjust. - periodSeconds: 10-15 seconds is a good balance between timely checks and resource usage; no need for 20s unless you have specific constraints.
- failureThreshold: For slow-starting apps, bump this to 5 to avoid unnecessary restarts.
- timeoutSeconds: Keep this short (3s is fine) unless your health check endpoint takes time to respond.
Fixes & Configuration Examples
Example 1: Spring Boot Actuator Health Checks
If you're using Spring Boot, leverage the actuator endpoints for granular liveness and readiness checks:
livenessProbe: httpGet: path: /actuator/health/liveness port: 8080 scheme: HTTP initialDelaySeconds: 60 periodSeconds: 10 timeoutSeconds: 3 failureThreshold: 3 readinessProbe: httpGet: path: /actuator/health/readiness port: 8080 scheme: HTTP initialDelaySeconds: 30 periodSeconds: 10 timeoutSeconds: 3 failureThreshold: 3
Note: Make sure to enable these endpoints in your Spring Boot application properties.
Example 2: Tomcat Manager Status Endpoint
Use this if your app doesn't have a custom health check:
livenessProbe: httpGet: path: /manager/status port: 8080 scheme: HTTP httpHeaders: - name: Authorization value: Basic cHJvYmUtdXNlcjpzZWN1cmUtcHJvYmUtcGFzcw== # Base64 of "probe-user:secure-probe-pass" initialDelaySeconds: 45 periodSeconds: 15 timeoutSeconds: 3 failureThreshold: 3 readinessProbe: httpGet: path: /manager/status port: 8080 scheme: HTTP httpHeaders: - name: Authorization value: Basic cHJvYmUtdXNlcjpzZWN1cmUtcHJvYmUtcGFzcw== initialDelaySeconds: 30 periodSeconds: 15 timeoutSeconds: 3 failureThreshold: 3
Tip: Generate the Base64 auth string with
echo -n "probe-user:secure-probe-pass" | base64.
Example 3: Adjust for Root-Deployed Apps
If your app is definitely deployed to Tomcat's ROOT context and / returns 200, fix the initial delay and failure threshold:
livenessProbe: failureThreshold: 5 httpGet: path: / port: 8080 scheme: HTTP initialDelaySeconds: 60 periodSeconds: 20 successThreshold: 1 timeoutSeconds: 3 readinessProbe: failureThreshold: 3 httpGet: path: / port: 8080 scheme: HTTP initialDelaySeconds: 40 periodSeconds: 15 successThreshold: 1 timeoutSeconds: 3
Troubleshooting Tips
- Test the probe path directly from inside the pod: Run
kubectl exec -it <pod-name> -- curl -v http://localhost:8080/<probe-path>to check the HTTP status code. - Check pod events:
kubectl describe pod <pod-name>will show exactly why the probe failed (e.g., connection timeout, non-200 status code).
内容的提问来源于stack exchange,提问作者Carsten

