如何删除已离线TiDB集群节点的监控数据?
Hey there! If you've got an offline node in your production TiDB cluster and want to clear out its old monitoring data, I've got you covered. TiDB's monitoring stack relies on Prometheus for time-series storage, so we'll focus on managing Prometheus data and tweaking related components to clean things up.
1. Identify the Offline Node's Unique Labels
First, you need to pinpoint exactly how Prometheus identifies the offline node. Head to your Prometheus UI (usually at http://<your-prometheus-ip>:9090/graph) and run this PromQL query to list all cluster targets:
up{job=~"tidb.*|pd.*|tikv.*"}
Look for entries where the up value is 0—that's your offline node. Note down labels like instance, job, or host that uniquely mark this node (e.g., instance="192.168.0.10:20180" for a TiKV node).
2. Delete the Target Node's Metrics via Prometheus API
Prometheus has an admin API for deleting time-series data. First, confirm your Prometheus instance has the web.enable-admin-api flag enabled (this is on by default in TiDB's standard deployment, but double-check if you've modified configs).
Use curl to send a DELETE request targeting the offline node. For example, to delete all metrics for the TiKV node we noted earlier:
curl -X DELETE http://<your-prometheus-ip>:9090/api/v1/admin/tsdb/delete_series \ -d 'match[]={instance="192.168.0.10:20180"}'
If you only want to delete data from a specific time range (say, before the node went offline), add start and end parameters using Unix timestamps:
curl -X DELETE http://<your-prometheus-ip>:9090/api/v1/admin/tsdb/delete_series \ -d 'match[]={instance="192.168.0.10:20180"}' \ -d 'start=1690000000' \ -d 'end=1700000000'
Pro tip: Test your match[] query in the Prometheus UI first to make sure it only targets the offline node—you don't want to accidentally delete data from active nodes!
3. Clean Up Prometheus Disk Space (Optional but Recommended)
Deleting series marks the data as "tombstoned" but doesn't immediately remove it from disk. To free up that space, run a compaction via the API:
curl -X POST http://<your-prometheus-ip>:9090/api/v1/admin/tsdb/clean_tombstones
This will permanently erase the tombstoned data and reclaim disk space.
4. Update Grafana Dashboards (If Needed)
If your Grafana dashboards still show the offline node in panels, you can adjust the queries to exclude it. For example, modify a TiKV metric query from:
sum(tikv_engine_apply_duration_seconds_sum{job="tikv"}) by (instance)
To:
sum(tikv_engine_apply_duration_seconds_sum{job="tikv", instance!="192.168.0.10:20180"}) by (instance)
Or a cleaner approach: filter using the up metric to only include active nodes:
sum(tikv_engine_apply_duration_seconds_sum{job="tikv", up="1"}) by (instance)
Key Notes
- Irreversible action: Once you delete the data, you can't get it back—make sure the node is permanently offline before proceeding.
- Managed TiDB services: If you're using TiDB Cloud or another managed service, you might not have direct access to the Prometheus API. Reach out to the service's support team for guidance on cleaning up offline node data.
- Broad matches: Avoid overly broad
match[]parameters (like justjob="tikv"), as this will delete data for all TiKV nodes, not just the offline one.
内容的提问来源于stack exchange,提问作者Caitin Chen

