关于Prometheus中label_replace函数及对应告警规则的功能咨询
Hey there! Let me break down this Prometheus alert rule for you clearly, no jargon overload promise 😊
First, let's look at when this alert will trigger — it fires if either of these two conditions is met:
- In the last 5 minutes, the number of changes to the
hikaricp_connections_timeout_totalmetric is greater than 2 - The
hikaricp_connectionsmetric doesn't exist at all (absent)
Now let's unpack each part one by one:
1. The trigger conditions explained
changes(hikaricp_connections_timeout_total[5m]) > 2hikaricp_connections_timeout_totalis a counter metric that tracks the total number of connection timeouts in your HikariCP connection pool (commonly used in Java applications).- The
changes()function counts how many times this counter's value has increased (since each timeout increments the counter). So this condition means "there have been more than 2 connection timeouts in the last 5 minutes" — a clear sign your app is struggling to establish or maintain database connections.
absent(hikaricp_connections) == 1hikaricp_connectionstracks the current number of active connections in the pool.- The
absent()function returns 1 only if the metric has no data at all. This condition triggers when Prometheus isn't receiving any connection pool metrics from the target instance — which could mean the app is down, the monitoring agent is broken, or the connection pool itself has crashed.
2. What the label_replace part does
The label_replace(..., "ip", "$1", "instance", "([^:]+)(:[0-9]+)?") part is just adding a practical label to the alert:
- It takes the existing
instancelabel (which usually looks like192.168.0.10:8080or just a raw IP) - Uses a regex to extract just the IP address (ignoring the port if one exists)
- Creates a new label called
ipwith that extracted value
This makes your alerts cleaner — instead of seeing a full instance string with a port, you just get the IP, which makes it way faster to identify which server or app is having issues.
To sum it up: This alert watches for two critical HikariCP connection pool problems — frequent timeouts (which hurt app performance) and complete loss of metrics (which signals a potential app/pool outage) — and adds an IP label to simplify troubleshooting.
备注:内容来源于stack exchange,提问作者Андрей

