如何正确使用Prometheus OpenMetrics回填?实测遇障求助
问题:Prometheus回填测试失败,重启后样本越界且查询无结果
环境说明
Win11主机,WSL2运行Ubuntu 24.04,部署minikube及cAdvisor DaemonSet;Prometheus 2.46.0以独立进程运行:
- 启动命令:
nohup ./prometheus --config.file=prometheus.yml --storage.tsdb.path=./data --storage.tsdb.retention.time=1y &
- 配置文件
prometheus.yml:
# my global config global: scrape_interval: 15s # Set the scrape interval to every 15 seconds. Default is every 1 minute. evaluation_interval: 15s # Evaluate rules every 15 seconds. The default is every 1 minute. # scrape_timeout is set to the global default (10s). # Alertmanager configuration alerting: alertmanagers: - static_configs: - targets: # - alertmanager:9093 # Load rules once and periodically evaluate them according to the global 'evaluation_interval'. rule_files: # - "first_rules.yml" # - "second_rules.yml" # A scrape configuration containing exactly one endpoint to scrape: # Here it's Prometheus itself. scrape_configs: # The job name is added as a label `job=<job_name>` to any timeseries scraped from this config. - job_name: "prometheus" # metrics_path defaults to '/metrics' # scheme defaults to 'http'. static_configs: - targets: ["localhost:9090"] - job_name: "cadvisor" static_configs: - targets: ["192.168.49.2:8080"]
正常启动时,Prometheus可正常采集两个端点指标,启停后运行稳定。
回填测试步骤
- 采集cAdvisor指标生成回填数据文件:
curl http://192.168.49.2:8080/metrics | grep "container_cpu_usage_seconds_total\|# ??? container_cpu_usage_seconds_total" | (head -n 2; tail -n 3) > backfill_test_data.prom echo "# EOF" >> backfill_test_data.prom
promtool check metrics提示标签命名格式警告,因正常采集无问题,暂忽略。
- 停止Prometheus并清理存储,生成回填块:
pkill prometheus rm -rf ./data ./promtool tsdb create-blocks-from openmetrics backfill_test_data.prom ./data
输出显示块生成成功:
BLOCK ULID MIN TIME MAX TIME DURATION NUM SAMPLES NUM CHUNKS NUM SERIES SIZE 01K959VCHJ6BS1BHZAXPS2E0BC 1762187220440000 1762187220966001 8m46.001s 3 3 3 10707
- 重启Prometheus后出现警告日志:
ts=2025-11-03T16:48:25.748Z caller=scrape.go:1729 level=warn component="scrape manager" scrape_pool=cadvisor target=http://192.168.49.2:8080/metrics msg="Error on ingesting samples that are too old or are too far into the future" num_dropped=2950 ts=2025-11-03T16:48:25.748Z caller=scrape.go:1338 level=warn component="scrape manager" scrape_pool=cadvisor target=http://192.168.49.2:8080/metrics msg="Appending scrape report failed" err="out of bounds" ts=2025-11-03T16:48:34.734Z caller=scrape.go:1729 level=warn component="scrape manager" scrape_pool=prometheus target=http://localhost:9090/metrics msg="Error on ingesting samples that are too old or are too far into the future" num_dropped=418 ts=2025-11-03T16:48:34.734Z caller=scrape.go:1338 level=warn component="scrape manager" scrape_pool=prometheus target=http://localhost:9090/metrics msg="Appending scrape report failed" err="out of bounds"
此时查询container_cpu_usage_seconds_total返回空结果,两个采集端点均出现样本越界问题。
问题排查与解决
核心原因
回填数据的时间戳(1762187220440000,对应2025年11月)与重启Prometheus的实时时间完全重叠,触发TSDB的时间边界冲突:
- Prometheus启动时会以当前时间作为存储的默认边界,当回填数据的时间范围与实时采集的时间范围重叠时,TSDB会拒绝写入新样本,同时可能导致回填的历史数据无法被正常检索。
- 直接截取实时指标生成的回填数据,时间戳与实时采集的新样本时间戳几乎一致,TSDB不允许同一时间范围的块重复存在,引发"out of bounds"错误。
修正步骤
1. 修改回填数据的时间戳
将回填数据的时间戳调整为早于当前时间的历史时段(比如提前1天),避免和实时时间重叠:
- 手动修改:打开
backfill_test_data.prom,将每行末尾的毫秒级时间戳减去86400000(1天的毫秒数),例如将1762187220440改为1762100820440。 - 脚本批量修改:
sed -i 's/\([0-9]\{13\}\)$/echo $((\1 - 86400000))/e' backfill_test_data.prom
2. 重新生成回填块
停止Prometheus,清理存储后重新执行回填:
pkill prometheus rm -rf ./data ./promtool tsdb create-blocks-from openmetrics backfill_test_data.prom ./data
3. 启动Prometheus(可选:指定时间边界)
为确保TSDB正确识别历史数据,启动时可手动指定存储的时间范围:
nohup ./prometheus --config.file=prometheus.yml --storage.tsdb.path=./data --storage.tsdb.retention.time=1y --storage.tsdb.min-time=1762100820000 --storage.tsdb.max-time=$(date +%s000) &
其中1762100820000是回填数据的最小时间戳(毫秒级),$(date +%s000)是当前时间的毫秒级时间戳。
4. 验证结果
重启后查看日志是否还有越界警告,在Prometheus GUI中查询container_cpu_usage_seconds_total,应能看到回填的历史数据,同时实时采集的新样本可正常写入。
额外注意事项
- 生成回填数据时,要确保时间戳属于合理的历史时段,避免与当前时间重叠或处于未来。
- 若回填数据时间戳在未来,Prometheus启动后会认为当前时间早于存储的最小时间,直接触发越界错误,无法处理实时数据。
内容的提问来源于stack exchange,提问作者scotofil
相关产品推荐
相关产品推荐

