如何通过Prometheus获取按应用、按天统计的变更前置时间(Lead time for changes)并集成至Grafana面板?
Do Prometheus have built-in metrics for lead time for changes?
Short answer: No, Prometheus doesn’t ship with out-of-the-box lead time metrics. This is because lead time depends entirely on your specific CI/CD pipeline and version control workflow—Prometheus only scrapes metrics you explicitly expose to it.
That said, you have two reliable paths to get this data into Prometheus:
1. Build a custom exporter
You can create a small script or lightweight service that pulls data from your Git repo (e.g., commit timestamps) and CI/CD tool (e.g., deployment completion times), calculates the time difference between code commit and production deployment, then exposes this as a Prometheus gauge metric.
A sample metric might look like this:
change_lead_time_seconds{application="payment-service",environment="prod",commit_hash="a1b2c3",day="2024-05-20"} 1620
- Include labels like
application,day, andenvironmentto make filtering and aggregation straightforward later. - The value represents the total seconds between when code was committed and when it finished deploying.
2. Use existing DORA metrics tools
There are open-source tools built specifically to export DORA metrics (including lead time) to Prometheus. These tools usually integrate with GitLab, GitHub, Jenkins, etc., out of the box—you just need to configure them to point to your CI/CD and Git systems, and they’ll handle calculating and exposing the metrics automatically.
Integrating the metric into Grafana
Once your lead time metric is being scraped by Prometheus, here’s how to build a useful, actionable Grafana dashboard:
1. Verify the metric exists first
Head to your Prometheus UI (typically http://<prometheus-ip>:9090) and run a query like change_lead_time_seconds to confirm the data is being pulled in correctly before moving to Grafana.
2. Set up your Grafana dashboard
- Add the Prometheus data source: If you haven’t already, go to Grafana’s Data Sources page, add Prometheus, and enter your Prometheus server URL.
- Create a new panel:
- Select Prometheus as the data source.
- Use a PromQL query to aggregate data by application and day. For example, to get the average lead time per app per day:
If you didn’t add aavg by (application, day) (change_lead_time_seconds)daylabel in your exporter, you can generate it on the fly using timestamp math:
(This converts the metric’s timestamp to the start of its corresponding day, so you can group results by date.)avg by (application, day_start) ( change_lead_time_seconds * on() group_left(day_start) label_replace( floor(timestamp(change_lead_time_seconds) / 86400) * 86400, "day_start", "$1", "", "(.*)" ) ) - For a more realistic view (avoiding skewed data from outliers), use quantiles instead of averages:
quantile(0.95, change_lead_time_seconds) by (application, day)
- Configure visualization:
- Use a line graph to track lead time trends over days, with the
applicationlabel as the legend to compare performance across apps. - Use a table panel to show a daily breakdown per app—this makes it easy to spot specific days with unusually long lead times.
- Use a line graph to track lead time trends over days, with the
- Add filters:
- Create a dashboard variable for
application(use the querylabel_values(change_lead_time_seconds, application)). This lets users filter the dashboard to view data for specific apps. - Add an
environmentvariable if you track lead time across staging, production, etc.
- Create a dashboard variable for
- Set refresh interval: Since this is daily data, set the dashboard refresh to once per day (or hourly if you want near-real-time updates for the current day’s deployments).
3. Save and share
Save the dashboard, and share it with your team to keep everyone aligned on deployment efficiency and identify bottlenecks early.
Pro Tips
- Ensure your CI/CD pipeline captures accurate timestamps: You need the exact time a commit was pushed and the exact time it finished deploying to production to calculate lead time correctly.
- Handle edge cases: Exclude rollbacks or emergency hotfixes if they skew your data, or add a
change_typelabel to filter them out in queries. - Combine with other DORA metrics: Add panels for deployment frequency, mean time to recover, and change failure rate to create a full, holistic DORA metrics dashboard.
内容的提问来源于stack exchange,提问作者MBA

