Prometheus标签中使用标识ID的最佳实践及可行性问询
Hey there! Let's tackle your two Prometheus questions head-on—they're super common when dealing with payment systems, so great asking them.
When working with IDs (like user IDs, order IDs) as Prometheus labels, the biggest thing to watch out for is cardinality—since high-cardinality labels can blow up your storage and slow down queries. Here are my go-to best practices:
- Only use IDs as labels if you need to slice/aggregate by them: If you never need to query "failures for user X" or "latency for order Y", don't add the ID as a label. Stick to low-cardinality dimensions like
service,region, orpayment_methodinstead. - Enforce consistent naming: Pick a convention (e.g.,
user_id,order_id) and stick to it across your team. This avoids confusion when writing queries or building dashboards. - Set strict retention for high-cardinality metrics: If you must use an ID label, configure a shorter retention period for those specific metrics. For example, you can use Prometheus's rule files to route high-cardinality data to a separate storage with lower retention, or use remote write to tools like Thanos/Cortex that handle high cardinality better.
- Test cardinality before production: Use a query like
count by (__name__, payment_id) ({payment_id=~".+"})in a staging environment to see how many unique time series you'll generate. If the number is in the tens of thousands or more, rethink your approach. - Avoid unbounded dynamic IDs: If the ID space is infinite (like payment IDs), think twice—each new ID creates a new time series, which is unsustainable long-term.
First, the straight answer: You can add payment_id as a label, but it's almost never a good idea. Each unique payment ID creates a new time series, which will quickly eat up your storage and make queries sluggish. But don't worry—there are better ways to get payment IDs into your failure alerts without killing your Prometheus instance.
Here are the most practical solutions:
- Log-based alerting (my top pick): Send all payment failure logs (including
payment_id) to a log system like Loki or ELK. Then set up alerts directly from the logs—tools like Alertmanager (with Loki integration) or Elastic Alerting can trigger alerts when a failure is detected, and include thepayment_idstraight from the log entry. This keeps Prometheus focused on high-level metrics, while logs handle the granular details. - Attach payment IDs as alert annotations (via exporter logic): If you want to stick with Prometheus alerts, modify your application or custom exporter to add
payment_idas an annotation when a failure occurs. For example:- Track a generic
payment_failure_totalcounter (nopayment_idlabel) - When a failure happens, have your exporter add metadata (including
payment_id) as an annotation to the metric - Alertmanager will pick up this annotation and display it in the alert notification, without creating a new time series for each payment ID
- Track a generic
- Enrich alerts via webhook: Trigger a Prometheus alert on the generic
payment_failure_totalmetric, then use an Alertmanager webhook to fetch the correspondingpayment_idfrom your database or log system. The webhook can add the ID to the alert before sending it to your team (e.g., via Slack, PagerDuty). - Selective sampling (last resort): If you absolutely need
payment_idin Prometheus, sample only a subset of failures (e.g., 10% of failed payments). This reduces cardinality while still giving you some IDs for debugging. Just note you won't get every failure's ID this way.
内容的提问来源于stack exchange,提问作者JosiDiez

