You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Prometheus Node Exporter排查连续X天宕机的服务器

Solution for Identifying Fully Unreachable Instances Over X Days

To find instances that have been completely unreachable for the past X days, use this combined Prometheus query:

absent_over_time(up{job="node_exporter"}[Xd]) == 1 OR max_over_time(up{job="node_exporter"}[Xd]) == 0

Breakdown of the Query:

  • Replace placeholders:

    • job="node_exporter": Adjust this to match the job label used for your Node Exporter instances.
    • Xd: Replace with your desired duration (e.g., 2d for 2 days, 7d for 7 days).
  • Two key conditions:

    1. absent_over_time(up[Xd]) == 1: Returns instances where Prometheus recorded no up metric samples at all in the X-day window. This covers servers that were completely unreachable (e.g., powered off, network disconnected) so Prometheus couldn’t even attempt to scrape them.
    2. max_over_time(up[Xd]) == 0: Returns instances where every scrape attempt in the X-day window returned up=0. This covers servers where Node Exporter was installed but consistently unresponsive during scrapes.

Why Your Original Query Didn’t Work:

min_over_time(up[2d]) == 0 checks if the minimum up value over 2 days is 0. This includes any instance that was down even once (e.g., a 1-minute outage), not just those that stayed down continuously. Using max_over_time ensures we only include instances where no successful scrapes occurred, and absent_over_time fills in the gap for instances that couldn’t be scraped at all.

内容的提问来源于stack exchange,提问作者KUMAR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 19:12:19