Prometheus告警咨询:僵尸进程与登录用户数指标及自定义扩展
Hey there! Let's break this down for you step by step—first, let's check if the metrics you need already exist, then cover how to add custom ones if not.
- Zombie Processes: You might have missed this, but node_exporter actually includes a built-in metric for zombie processes:
node_procs_zombie. This metric tracks the exact count of zombie processes on the node, so you can use it directly for your alerting rules without any custom code. - Logged-in User Count: Unfortunately, node_exporter doesn't provide a default metric for this. We'll need to build a custom solution to collect this data and feed it into node_exporter.
The easiest way to add custom metrics is using node_exporter's textfile collector. This feature lets you write custom metrics to a plain text file, and node_exporter will automatically scrape and expose those metrics alongside its default ones. No need to build a separate exporter service!
Here's how to implement this with Shell, Python, or Go:
2.1 Shell Script Implementation
Create a simple shell script to calculate logged-in users and write the metric to the textfile directory:
#!/bin/bash # Define the output directory (match this with node_exporter's config) TEXT_DIR="/var/lib/node_exporter/textfile_collector" mkdir -p "$TEXT_DIR" # Calculate logged-in user count using `who` command USER_COUNT=$(who | wc -l) # Write the metric to a .prom file echo "node_logged_in_users $USER_COUNT" > "$TEXT_DIR/login_users.prom"
- Make the script executable:
chmod +x /path/to/login_users.sh - Set up a cron job to run it every minute (ensure metrics stay up-to-date):
* * * * * root /path/to/login_users.sh - Update your node_exporter startup command to include the textfile collector flag:
node_exporter --collector.textfile.directory=/var/lib/node_exporter/textfile_collector
2.2 Python Script Implementation
If you prefer Python, here's a more robust script:
#!/usr/bin/env python3 import subprocess import os def get_logged_in_users(): # Use `who -q` to get a concise user count output result = subprocess.run(["who", "-q"], capture_output=True, text=True) for line in result.stdout.splitlines(): if line.startswith("# users="): return int(line.split("=")[1]) return 0 def main(): text_dir = "/var/lib/node_exporter/textfile_collector" os.makedirs(text_dir, exist_ok=True) user_count = get_logged_in_users() with open(f"{text_dir}/login_users.prom", "w") as f: f.write(f"node_logged_in_users {user_count}\n") if __name__ == "__main__": main()
- Make it executable:
chmod +x /path/to/login_users.py - Set up a cron job to run it every minute, just like the shell script.
2.3 Go Implementation (For Higher Performance)
If you need a more efficient, compiled solution, use Go:
package main import ( "fmt" "os" "os/exec" "strings" ) func getLoggedInUsers() (int, error) { cmd := exec.Command("who", "-q") output, err := cmd.Output() if err != nil { return 0, err } for _, line := range strings.Split(string(output), "\n") { if strings.HasPrefix(line, "# users=") { var count int _, err := fmt.Sscanf(line, "# users=%d", &count) if err != nil { return 0, err } return count, nil } } return 0, nil } func main() { count, err := getLoggedInUsers() if err != nil { fmt.Printf("Failed to get user count: %v\n", err) os.Exit(1) } textDir := "/var/lib/node_exporter/textfile_collector" if err := os.MkdirAll(textDir, 0755); err != nil { fmt.Printf("Failed to create directory: %v\n", err) os.Exit(1) } file, err := os.Create(fmt.Sprintf("%s/login_users.prom", textDir)) if err != nil { fmt.Printf("Failed to create metric file: %v\n", err) os.Exit(1) } defer file.Close() if _, err := file.WriteString(fmt.Sprintf("node_logged_in_users %d\n", count)); err != nil { fmt.Printf("Failed to write metric: %v\n", err) os.Exit(1) } }
- Compile the binary:
go build -o login_users_exporter login_users.go - Run the binary via cron or a systemd timer to keep metrics fresh.
Now that you have both metrics, here's a sample alert rule file to monitor them:
groups: - name: system_health rules: # Alert for high zombie process count - alert: HighZombieProcesses expr: node_procs_zombie > 5 for: 1m labels: severity: warning annotations: summary: "High zombie process count on {{ $labels.instance }}" description: "{{ $labels.instance }} has {{ $value }} zombie processes (threshold: 5)." # Alert for excessive logged-in users - alert: TooManyLoggedInUsers expr: node_logged_in_users > 10 for: 1m labels: severity: warning annotations: summary: "Too many logged-in users on {{ $labels.instance }}" description: "{{ $labels.instance }} has {{ $value }} active logged-in users (threshold: 10)."
Add this file to your Prometheus rule_files configuration to start receiving alerts.
内容的提问来源于stack exchange,提问作者anujkum

