You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Telegraf监控Nginx:请求处理分布与500错误排查需求

Solution: Track Nginx Request Handlers & Status Codes Without Overly Complex Logging

Hey there, let's tackle this problem step by step—you’re right that out-of-the-box tools like Telegraf’s Nginx plugin don’t cover all these granular needs, but there are practical, lightweight ways to get the metrics you want without spinning up a full ELK stack or similar heavy logging infrastructure.

1. Identifying & Tracking Request Handler Distribution

The core challenge here is tagging each request with its handler (Rails, Nginx static, nginx_status) so you can aggregate counts. Here are two reliable approaches:

Option 1: Custom Nginx Log Fields (No Extra Dependencies)

First, define a mapping in your Nginx config to tag each request with its handler. Add this to your http block:

# Map request paths to handler tags
map $request_uri $handler {
    ~^/nginx_status$ nginx_status;
    ~*\.(css|js|png|jpg|svg)$ nginx_static; # Nginx directly serves these
    default rails; # All other requests proxy to Rails
}

# Update your log format to include the handler tag
log_format monitor '$remote_addr - $remote_user [$time_local] "$request" '
                  '$status $body_bytes_sent "$http_referer" '
                  '"$http_user_agent" "$handler"';

# Apply this log format to your server/location blocks
access_log /var/log/nginx/monitor_access.log monitor;

This creates a dedicated log file with only the fields you need (including $handler). Then configure Telegraf’s tail plugin to parse this log with a grok pattern:

[[inputs.tail]]
  files = ["/var/log/nginx/monitor_access.log"]
  from_beginning = false
  pipe = false
  grok_patterns = ['%{COMBINED_LOG_FORMAT} %{QS:handler}']
  tag_keys = ["handler", "status"]

Telegraf will send metrics tagged with handler and status to your time-series database (InfluxDB/Prometheus), where you can easily calculate the percentage of requests per handler.

Option 2: OpenResty Lua Scripting (Real-Time Metrics)

If you’re using OpenResty (Nginx + Lua), you can track counts in-memory without relying on logs. Add this to your Nginx config:

http {
    # Create a shared memory zone for metrics
    lua_shared_dict metrics_dict 10m;

    # Track handler and status code on request finish
    log_by_lua_block {
        local handler = ngx.var.handler # Reuse the $handler map from Option 1
        local status = tostring(ngx.status)
        local dict = ngx.shared.metrics_dict
        
        # Increment counters for handler + status
        dict:incr(handler .. "_total", 1, 0)
        dict:incr(handler .. "_status_" .. status, 1, 0)
    }

    # Expose a custom metrics endpoint for Telegraf to scrape
    location /custom_nginx_metrics {
        access_log off;
        allow 127.0.0.1; # Restrict to Telegraf's IP
        deny all;
        content_by_lua_block {
            local dict = ngx.shared.metrics_dict
            local keys = dict:get_keys(0)
            ngx.say("# HELP nginx_requests_total Total requests per handler")
            ngx.say("# TYPE nginx_requests_total counter")
            for _, key in ipairs(keys) do
                if string.match(key, "_total$") then
                    local handler = string.gsub(key, "_total$", "")
                    ngx.say("nginx_requests_total{handler=\"" .. handler .. "\"} " .. dict:get(key))
                end
            end
            # Add similar output for status codes
        }
    }
}

Then configure Telegraf’s http plugin to scrape /custom_nginx_metrics periodically—this gives you real-time, log-free metrics.

2. Tracking HTTP Status Codes & 500 Error Peaks

Once you have handler-tagged metrics, tracking status codes (including 500 peaks) becomes straightforward:

  • In Grafana/Chronograf: Build dashboards that show:
    • Total requests per status code, filtered by handler
    • 1-minute/5-minute rolling averages for 500 errors
  • Alerting: Set up threshold alerts (e.g., "If 500 errors from Rails exceed 10 in 1 minute, trigger an alert"). Most monitoring tools (Prometheus Alertmanager, InfluxDB Alerts) support this natively.
  • Peak Analysis: Use your time-series database to query historical 500 error peaks alongside handler tags—this lets you quickly pinpoint if failures are coming from Rails, Nginx, or another source.

3. How This Solves Your Core Goals

  • Capacity Planning: By tracking long-term growth in rails handler requests (combined with Rails response time metrics), you can clearly identify when to scale your Rails application servers. Similarly, growth in nginx_static requests might indicate a need to optimize static resource caching or scale Nginx.
  • Fault Detection: Real-time alerts on 500 peaks (tagged by handler) let you triage issues faster—if the peak is from Rails, you can jump straight to Rails logs; if it’s from Nginx, check for config errors or static resource issues.

内容的提问来源于stack exchange,提问作者jma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:54:00