Prometheus Python客户端导入历史性能数据时间戳超出范围报错
Prometheus 历史性能测试结果导入异常问题
问题背景
尝试将性能测试历史结果导入Prometheus时,使用官方Python Prometheus客户端出现指标无法入库的异常。
可正常运行的代码示例
dt_now = datetime.datetime.now(tz=pytz.timezone('UTC')) gobj = GaugeMetricFamily('FooMetricGood', '') gobj.add_metric([], 123, timestamp=dt_now.timestamp()) yield gobj
异常代码示例
dt_format = '%Y-%m-%d_%H-%M-%S.%f %z' dt_custom_str = '2021-11-11_18-12-59.000000 +0000' dt_parsed_from_custom = datetime.datetime.strptime(dt_custom_str, dt_format) gobj = GaugeMetricFamily('FooMetricNotWorking', '') gobj.add_metric([], 789987, timestamp=dt_parsed_from_custom.timestamp()) yield gobj
首次告警日志
prometheus-prometheus-1 | ts=2021-11-11T13:41:01.895Z caller=scrape.go:1563 level=warn component="scrape manager" scrape_pool=services target=http://192.168.64.1:8080/metrics msg="Error on ingesting samples that are too old or are too far into the future" num_dropped=1
补充验证信息
客户端指标输出(正常展示)
# HELP FooMetricGood # TYPE FooMetricGood gauge FooMetricGood 123.0 1639475119451 # HELP FooMetricGoodToo # TYPE FooMetricGoodToo gauge FooMetricGoodToo 456.0 1639475119451 # HELP FooMetricNotWorkingNew # TYPE FooMetricNotWorkingNew gauge FooMetricNotWorkingNew 789987.0 1639355699000
服务端最新报错日志
prometheus-prometheus-1 | ts=2021-12-14T09:51:35.524Z caller=scrape.go:1611 level=debug component="scrape manager" scrape_pool=services target=http://192.168.64.1:8080/metrics msg="Out of bounds metric" series=FooMetricNotWorkingNew prometheus-prometheus-1 | ts=2021-12-14T09:51:35.524Z caller=scrape.go:1563 level=warn component="scrape manager" scrape_pool=services target=http://192.168.64.1:8080/metrics msg="Error on ingesting samples that are too old or are too far into the future" num_dropped=1
时间戳验证结果
- 正常指标时间戳:
1639475119451 - 异常指标时间戳:
1639355699000 - 二者时间差为119420451毫秒,折合33.17小时
- 调整自定义时间与当前时间差为1.29小时后,问题仍复现
问题原因与解决方案
根因分析
- Prometheus默认开启样本时间戳容忍窗口机制,启动参数
--scrape.sample-tolerance默认值为5分钟,只要样本时间戳与服务端当前时间差超过该阈值,无论时间是否合法都会被判定为越界丢弃,该设计是为了避免时序数据紊乱。 - 代码中时间戳传参存在单位错误:Python
datetime.timestamp()方法返回的是秒级浮点数值,而Prometheus Python客户端add_metric方法的timestamp参数要求传入毫秒级整数,直接传入秒级数值会导致时间戳解析异常,进一步触发越界判定。
修复方案
临时修复(少量历史数据导入)
- 调整Prometheus启动参数,放大时间戳容忍窗口到覆盖历史数据跨度,例如要导入72小时内的历史数据,新增启动参数:
--scrape.sample-tolerance=72h - 修正代码中的时间戳传参逻辑,转换为毫秒级整数:
gobj.add_metric([], 789987, timestamp=int(dt_parsed_from_custom.timestamp() * 1000))
最优方案(大量历史数据批量导入)
如果需要导入大量历史性能测试结果,不建议走常规抓取链路,推荐使用Prometheus官方远程写API或者promtool工具批量导入,无需修改服务端默认配置,也避免了单次抓取导入的性能限制。
内容的提问来源于stack exchange,提问作者Dmitrii Vinokurov
相关产品推荐
相关产品推荐

