如何针对计数器值增长设置10分钟间隔的告警(基于mtail监控PHP-FPM错误日志场景)
Prometheus Alert Rule for mtail's PHP-FPM Error Monitoring
Let's confirm and break down the exact configuration you need to avoid spamming alerts for every single error, and instead only alert when new errors occur within a 10-minute window.
Exact Alert Rule Configuration
Here's the complete, valid Prometheus alert rule snippet tailored to your needs:
groups: - name: php_fpm_errors rules: - alert: PHPFPMNewErrorsDetected expr: increase(php_fpm_errors_total[10m]) > 0 for: 10m labels: severity: warning annotations: summary: "New PHP-FPM errors detected in the last 10 minutes" description: "{{ $value }} new PHP-FPM errors were logged in the past 10 minutes."
Detailed Explanation
Let's walk through why each part works for your use case:
expr: increase(php_fpm_errors_total[10m]) > 0php_fpm_errors_totalis a counter metric generated by mtail, which increments every time a new error line is found in your logs. Counters only go up (or reset to 0 on process restart), so we use theincrease()function instead of looking at the raw value.increase(metric[10m])calculates the total increment of the counter over the past 10 minutes. Setting this to> 0means we're checking if any new errors were added in that window. This prevents alerts for every individual error—instead, we're checking for activity over the entire 10-minute period.
for: 10m- This parameter tells Prometheus to wait until the expression evaluates to
truefor a full 10 minutes before firing the alert. To align with your goal:- This setting ensures you won't get repeated alerts within the same 10-minute window. It will trigger once the window of error activity has completed, and stay resolved until the next 10-minute window with new errors.
- If you wanted to alert immediately when an error is detected but still avoid spam, you could remove the
forparameter—but the10mvalue here is better for reducing noise by only alerting after confirming sustained error activity over the window.
- This parameter tells Prometheus to wait until the expression evaluates to
Key Background Context
- Mtail's counter metrics are perfect for this use case because they naturally accumulate error counts over time. Using
increase()on a counter is the standard Prometheus practice for measuring growth over a time window, as it handles counter resets gracefully. - The alert rule's logic ensures you don't get flooded with alerts for every single error line. Instead, you get a single alert per 10-minute window where new errors were logged.
内容的提问来源于stack exchange,提问作者berkut
相关产品推荐
相关产品推荐

