Varnish Cache v4后端健康检查异常:CMS宕机仍返回正常状态
Hey there! Let's figure out why Varnish isn't recognizing your CMS backend as down when you shut it down manually—this is a common gotcha with Varnish's default health checking behavior.
1. The Core Issue: Missing Health Probe Configuration
By default, Varnish 4 doesn't perform active health checks on backends unless you explicitly define a probe. Without a probe, std.health(req.backend_hint) will always return true, even if your CMS is completely offline. Varnish assumes the backend is healthy unless told otherwise via a probe or manual CLI intervention.
2. Fix: Add a Backend Health Probe
You need to define a probe in your VCL to let Varnish actively check if your CMS is responding. Here's a sample configuration tailored to your use case:
backend cms { .host = "your-cms-hostname-or-ip"; .port = "80"; # or 443 if using HTTPS (adjust accordingly) # Define health check probe .probe = { .url = "/"; # Use a lightweight endpoint like /health if your CMS has one .interval = 5s; # Check every 5 seconds .timeout = 2s; # Fail if no response within 2 seconds .window = 5; # Evaluate the last 5 check attempts .threshold = 3; # Mark backend as sick if 3+ checks fail } }
- Use a dedicated health endpoint (like
/health) if your CMS provides one—it's more reliable than the homepage, which might have heavy assets. - Adjust
interval,timeout,window, andthresholdbased on your tolerance for downtime detection speed.
3. Ensure Grace Mode is Properly Configured
Even with a working probe, you need to make sure Varnish uses cached content during the grace period when the backend is down. Add these sections to your VCL:
Set Grace Period on Backend Responses
sub vcl_backend_response { # Set how long Varnish can use stale cached content when backend is down set beresp.grace = 24h; # Adjust to your desired grace period }
Use Grace Content When Backend is Unhealthy
sub vcl_recv { # Tell Varnish to allow grace content if backend is sick if (!std.health(req.backend_hint)) { set req.grace = 24h; # Match the beresp.grace value } } sub vcl_hit { # Deliver stale content from grace period if backend is down if (!std.health(req.backend_hint)) { if (obj.ttl + obj.grace > 0s) { return (deliver); } } }
4. Verify the Fix
After updating your VCL and reloading Varnish (varnishreload), test it out:
- Shut down your CMS manually.
- Wait for the probe's detection window (e.g., 15 seconds if using the sample probe settings).
- Check the backend health status via the Varnish CLI:
You should see yourvarnishadm backend.listcmsbackend marked assickinstead ofhealthy (no probe). - Send a request to Varnish—you should get the cached content instead of a 503.
Quick Notes
- If you're using HTTPS for your backend, make sure you're using
backend_ssl(or appropriate SSL configuration for Varnish 4) and adjust the probe URL/port accordingly. - Avoid setting the probe interval too low (e.g., <1s) as it can add unnecessary load to your CMS.
Give these steps a shot, and you should see Varnish correctly detect the down backend and serve cached grace content as expected!
内容的提问来源于stack exchange,提问作者David Beaudway

