基于Telegraf监控Nginx:请求处理分布与500错误排查需求
Hey there, let's tackle this problem step by step—you’re right that out-of-the-box tools like Telegraf’s Nginx plugin don’t cover all these granular needs, but there are practical, lightweight ways to get the metrics you want without spinning up a full ELK stack or similar heavy logging infrastructure.
1. Identifying & Tracking Request Handler Distribution
The core challenge here is tagging each request with its handler (Rails, Nginx static, nginx_status) so you can aggregate counts. Here are two reliable approaches:
Option 1: Custom Nginx Log Fields (No Extra Dependencies)
First, define a mapping in your Nginx config to tag each request with its handler. Add this to your http block:
# Map request paths to handler tags map $request_uri $handler { ~^/nginx_status$ nginx_status; ~*\.(css|js|png|jpg|svg)$ nginx_static; # Nginx directly serves these default rails; # All other requests proxy to Rails } # Update your log format to include the handler tag log_format monitor '$remote_addr - $remote_user [$time_local] "$request" ' '$status $body_bytes_sent "$http_referer" ' '"$http_user_agent" "$handler"'; # Apply this log format to your server/location blocks access_log /var/log/nginx/monitor_access.log monitor;
This creates a dedicated log file with only the fields you need (including $handler). Then configure Telegraf’s tail plugin to parse this log with a grok pattern:
[[inputs.tail]] files = ["/var/log/nginx/monitor_access.log"] from_beginning = false pipe = false grok_patterns = ['%{COMBINED_LOG_FORMAT} %{QS:handler}'] tag_keys = ["handler", "status"]
Telegraf will send metrics tagged with handler and status to your time-series database (InfluxDB/Prometheus), where you can easily calculate the percentage of requests per handler.
Option 2: OpenResty Lua Scripting (Real-Time Metrics)
If you’re using OpenResty (Nginx + Lua), you can track counts in-memory without relying on logs. Add this to your Nginx config:
http { # Create a shared memory zone for metrics lua_shared_dict metrics_dict 10m; # Track handler and status code on request finish log_by_lua_block { local handler = ngx.var.handler # Reuse the $handler map from Option 1 local status = tostring(ngx.status) local dict = ngx.shared.metrics_dict # Increment counters for handler + status dict:incr(handler .. "_total", 1, 0) dict:incr(handler .. "_status_" .. status, 1, 0) } # Expose a custom metrics endpoint for Telegraf to scrape location /custom_nginx_metrics { access_log off; allow 127.0.0.1; # Restrict to Telegraf's IP deny all; content_by_lua_block { local dict = ngx.shared.metrics_dict local keys = dict:get_keys(0) ngx.say("# HELP nginx_requests_total Total requests per handler") ngx.say("# TYPE nginx_requests_total counter") for _, key in ipairs(keys) do if string.match(key, "_total$") then local handler = string.gsub(key, "_total$", "") ngx.say("nginx_requests_total{handler=\"" .. handler .. "\"} " .. dict:get(key)) end end # Add similar output for status codes } } }
Then configure Telegraf’s http plugin to scrape /custom_nginx_metrics periodically—this gives you real-time, log-free metrics.
2. Tracking HTTP Status Codes & 500 Error Peaks
Once you have handler-tagged metrics, tracking status codes (including 500 peaks) becomes straightforward:
- In Grafana/Chronograf: Build dashboards that show:
- Total requests per status code, filtered by handler
- 1-minute/5-minute rolling averages for 500 errors
- Alerting: Set up threshold alerts (e.g., "If 500 errors from Rails exceed 10 in 1 minute, trigger an alert"). Most monitoring tools (Prometheus Alertmanager, InfluxDB Alerts) support this natively.
- Peak Analysis: Use your time-series database to query historical 500 error peaks alongside handler tags—this lets you quickly pinpoint if failures are coming from Rails, Nginx, or another source.
3. How This Solves Your Core Goals
- Capacity Planning: By tracking long-term growth in
railshandler requests (combined with Rails response time metrics), you can clearly identify when to scale your Rails application servers. Similarly, growth innginx_staticrequests might indicate a need to optimize static resource caching or scale Nginx. - Fault Detection: Real-time alerts on 500 peaks (tagged by handler) let you triage issues faster—if the peak is from Rails, you can jump straight to Rails logs; if it’s from Nginx, check for config errors or static resource issues.
内容的提问来源于stack exchange,提问作者jma

