You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Telegraf 1小时间隔指标丢失:如何延长指标保留时长?

解决Telegraf采集指标在Prometheus中快速丢失的问题

针对你遇到的「Telegraf每小时采集的指标10-15分钟后丢失,Grafana出现无数据空白」的问题,可以从以下几个方向排查和调整:

1. 调整Prometheus的指标保留策略

Prometheus默认保留指标15天,但如果配置了过短的保留时长或存储大小限制,可能导致指标被提前清理。修改prometheus.yml中的存储配置:

global:
  scrape_interval: 1h  # 和Telegraf的采集间隔对齐,避免不必要的频繁抓取
  evaluation_interval: 1h

storage:
  tsdb:
    retention.time: 30d  # 根据需求延长,比如设为30天
    # 如果配置了存储大小限制,可注释或调大
    # retention.size: 10GB

2. 确保Telegraf输出时保留原始时间戳

如果Telegraf输出的指标没有携带采集时的原始时间戳,Prometheus会用抓取时间作为时间戳,可能导致旧数据被误判为过期。在Telegraf的Prometheus输出插件中开启时间戳保留:
比如使用prometheus_client输出时:

[[outputs.prometheus_client]]
  listen = ":9273"
  add_timestamp = true  # 强制保留采集到的原始时间戳

如果用远程写入到Prometheus,同样要确保输出配置中没有丢弃时间戳。

3. 让Prometheus尊重指标的原始时间戳

默认情况下Prometheus可能会忽略指标自带的时间戳,改用抓取时间,这会导致间隔1小时的指标被集中到抓取时间点,后续被清理。在Prometheus的抓取配置中开启honor_timestamps:

scrape_configs:
  - job_name: 'telegraf'
    scrape_interval: 1h
    scrape_timeout: 5m
    honor_timestamps: true  # 启用后,Prometheus会使用指标本身的时间戳
    static_configs:
      - targets: ['telegraf:9273']

4. 排查Telegraf是否存在丢数情况

有时候指标丢失不是Prometheus的问题,而是Telegraf采集或输出环节出了问题。开启Telegraf的调试日志检查:

[agent]
  debug = true
  logfile = "/var/log/telegraf/telegraf.log"

查看日志中是否有输出失败、指标被过滤(比如标签不符合Prometheus规范)的记录,针对性调整Telegraf的处理器插件(如filter、rename)修正问题。

5. 用Prometheus记录规则填充数据空白(兜底方案)

如果偶尔出现采集中断导致的空白,可以通过记录规则用历史值填充缺失时段。在Prometheus中添加规则:

groups:
  - name: fill_missing_data
    rules:
      - record: your_metric:filled
        expr: last_over_time(your_metric[1h]) or vector(0)
        # 用最近1小时内的最后一个有效值填充,无值时用0(可根据业务调整)

之后在Grafana中使用这个填充后的指标,避免表格出现空白。

内容的提问来源于stack exchange,提问作者Grimzly

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 16:55:21