如何用Prometheus Node Exporter排查连续X天宕机的服务器
Solution for Identifying Fully Unreachable Instances Over X Days
To find instances that have been completely unreachable for the past X days, use this combined Prometheus query:
absent_over_time(up{job="node_exporter"}[Xd]) == 1 OR max_over_time(up{job="node_exporter"}[Xd]) == 0
Breakdown of the Query:
Replace placeholders:
job="node_exporter": Adjust this to match the job label used for your Node Exporter instances.Xd: Replace with your desired duration (e.g.,2dfor 2 days,7dfor 7 days).
Two key conditions:
absent_over_time(up[Xd]) == 1: Returns instances where Prometheus recorded noupmetric samples at all in the X-day window. This covers servers that were completely unreachable (e.g., powered off, network disconnected) so Prometheus couldn’t even attempt to scrape them.max_over_time(up[Xd]) == 0: Returns instances where every scrape attempt in the X-day window returnedup=0. This covers servers where Node Exporter was installed but consistently unresponsive during scrapes.
Why Your Original Query Didn’t Work:
min_over_time(up[2d]) == 0 checks if the minimum up value over 2 days is 0. This includes any instance that was down even once (e.g., a 1-minute outage), not just those that stayed down continuously. Using max_over_time ensures we only include instances where no successful scrapes occurred, and absent_over_time fills in the gap for instances that couldn’t be scraped at all.
内容的提问来源于stack exchange,提问作者KUMAR
相关产品推荐
相关产品推荐

