You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Prometheus告警咨询:僵尸进程与登录用户数指标及自定义扩展

Hey there! Let's break this down for you step by step—first, let's check if the metrics you need already exist, then cover how to add custom ones if not.

1. Check if Default Metrics Exist
  • Zombie Processes: You might have missed this, but node_exporter actually includes a built-in metric for zombie processes: node_procs_zombie. This metric tracks the exact count of zombie processes on the node, so you can use it directly for your alerting rules without any custom code.
  • Logged-in User Count: Unfortunately, node_exporter doesn't provide a default metric for this. We'll need to build a custom solution to collect this data and feed it into node_exporter.
2. Add Custom Metrics to Node Exporter

The easiest way to add custom metrics is using node_exporter's textfile collector. This feature lets you write custom metrics to a plain text file, and node_exporter will automatically scrape and expose those metrics alongside its default ones. No need to build a separate exporter service!

Here's how to implement this with Shell, Python, or Go:

2.1 Shell Script Implementation

Create a simple shell script to calculate logged-in users and write the metric to the textfile directory:

#!/bin/bash
# Define the output directory (match this with node_exporter's config)
TEXT_DIR="/var/lib/node_exporter/textfile_collector"
mkdir -p "$TEXT_DIR"

# Calculate logged-in user count using `who` command
USER_COUNT=$(who | wc -l)

# Write the metric to a .prom file
echo "node_logged_in_users $USER_COUNT" > "$TEXT_DIR/login_users.prom"
  • Make the script executable: chmod +x /path/to/login_users.sh
  • Set up a cron job to run it every minute (ensure metrics stay up-to-date):
    * * * * * root /path/to/login_users.sh
    
  • Update your node_exporter startup command to include the textfile collector flag:
    node_exporter --collector.textfile.directory=/var/lib/node_exporter/textfile_collector
    

2.2 Python Script Implementation

If you prefer Python, here's a more robust script:

#!/usr/bin/env python3
import subprocess
import os

def get_logged_in_users():
    # Use `who -q` to get a concise user count output
    result = subprocess.run(["who", "-q"], capture_output=True, text=True)
    for line in result.stdout.splitlines():
        if line.startswith("# users="):
            return int(line.split("=")[1])
    return 0

def main():
    text_dir = "/var/lib/node_exporter/textfile_collector"
    os.makedirs(text_dir, exist_ok=True)
    
    user_count = get_logged_in_users()
    with open(f"{text_dir}/login_users.prom", "w") as f:
        f.write(f"node_logged_in_users {user_count}\n")

if __name__ == "__main__":
    main()
  • Make it executable: chmod +x /path/to/login_users.py
  • Set up a cron job to run it every minute, just like the shell script.

2.3 Go Implementation (For Higher Performance)

If you need a more efficient, compiled solution, use Go:

package main

import (
	"fmt"
	"os"
	"os/exec"
	"strings"
)

func getLoggedInUsers() (int, error) {
	cmd := exec.Command("who", "-q")
	output, err := cmd.Output()
	if err != nil {
		return 0, err
	}
	for _, line := range strings.Split(string(output), "\n") {
		if strings.HasPrefix(line, "# users=") {
			var count int
			_, err := fmt.Sscanf(line, "# users=%d", &count)
			if err != nil {
				return 0, err
			}
			return count, nil
		}
	}
	return 0, nil
}

func main() {
	count, err := getLoggedInUsers()
	if err != nil {
		fmt.Printf("Failed to get user count: %v\n", err)
		os.Exit(1)
	}

	textDir := "/var/lib/node_exporter/textfile_collector"
	if err := os.MkdirAll(textDir, 0755); err != nil {
		fmt.Printf("Failed to create directory: %v\n", err)
		os.Exit(1)
	}

	file, err := os.Create(fmt.Sprintf("%s/login_users.prom", textDir))
	if err != nil {
		fmt.Printf("Failed to create metric file: %v\n", err)
		os.Exit(1)
	}
	defer file.Close()

	if _, err := file.WriteString(fmt.Sprintf("node_logged_in_users %d\n", count)); err != nil {
		fmt.Printf("Failed to write metric: %v\n", err)
		os.Exit(1)
	}
}
  • Compile the binary: go build -o login_users_exporter login_users.go
  • Run the binary via cron or a systemd timer to keep metrics fresh.
3. Prometheus Alerting Rules

Now that you have both metrics, here's a sample alert rule file to monitor them:

groups:
- name: system_health
  rules:
  # Alert for high zombie process count
  - alert: HighZombieProcesses
    expr: node_procs_zombie > 5
    for: 1m
    labels:
      severity: warning
    annotations:
      summary: "High zombie process count on {{ $labels.instance }}"
      description: "{{ $labels.instance }} has {{ $value }} zombie processes (threshold: 5)."
  # Alert for excessive logged-in users
  - alert: TooManyLoggedInUsers
    expr: node_logged_in_users > 10
    for: 1m
    labels:
      severity: warning
    annotations:
      summary: "Too many logged-in users on {{ $labels.instance }}"
      description: "{{ $labels.instance }} has {{ $value }} active logged-in users (threshold: 10)."

Add this file to your Prometheus rule_files configuration to start receiving alerts.

内容的提问来源于stack exchange,提问作者anujkum

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:09:13