You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

自定义Exporter处理时长超Prometheus抓取间隔时无指标返回

问题根源

Prometheus默认抓取超时为10秒,当你的collect()方法执行时间超过这个阈值时,Prometheus会主动终止请求,导致无法获取任何指标,因此图表呈现空白。你的代码中串行执行3次check_stream,每次耗时5秒,总时长15秒,超过了默认超时限制。


解决方案

方案1:并行执行采集任务,缩短总耗时

将串行的check_stream调用改为并行执行,利用线程池减少整体耗时,确保总时长低于Prometheus超时阈值。

修改后的代码:

#!/usr/bin/env python3

from prometheus_client import start_http_server, REGISTRY
from prometheus_client.metrics_core import GaugeMetricFamily
import time
from concurrent.futures import ThreadPoolExecutor

count = 0
# 初始化线程池,最多同时运行3个线程
executor = ThreadPoolExecutor(max_workers=3)

class TestExporter:
    
    def collect(self):
        global count
        count += 5
        
        # 并行提交3个采集任务
        futures = [executor.submit(self.check_stream, count, i) for i in range(3)]
        # 依次获取任务结果并返回
        for future in futures:
            metric = future.result()
            yield metric
    
    def check_stream(self, count, host):
        time.sleep(5)
        metric = GaugeMetricFamily('aa_stream_test', 'testing stream delete', labels=['stream'])
        if count > 100:
            metric.add_metric(['B', str(host)], count)
        else:
            metric.add_metric(['A', str(host)], count)
        print(count, 'registering . . .')
        return metric

if __name__ == '__main__':
    REGISTRY.register(TestExporter())
    start_http_server(8000)
    print("started server...")
    
    while True:
        time.sleep(1)

方案2:异步预采集+缓存,快速响应抓取请求

如果单任务耗时过长,并行仍无法满足超时要求,可以在后台线程定期采集指标并缓存,collect()方法直接返回缓存结果,确保响应时间极短。

修改后的代码:

#!/usr/bin/env python3

from prometheus_client import start_http_server, REGISTRY
from prometheus_client.metrics_core import GaugeMetricFamily
import time
import threading

count = 0
cached_metrics = []
# 线程锁保证缓存更新的线程安全
lock = threading.Lock()

def background_collect():
    global count, cached_metrics
    while True:
        start_time = time.time()
        new_count = count + 5
        new_metrics = []
        
        # 并行采集指标(也可根据需求串行)
        for i in range(3):
            time.sleep(5)
            metric = GaugeMetricFamily('aa_stream_test', 'testing stream delete', labels=['stream'])
            if new_count > 100:
                metric.add_metric(['B', str(i)], new_count)
            else:
                metric.add_metric(['A', str(i)], new_count)
            new_metrics.append(metric)
        
        # 加锁更新缓存
        with lock:
            count = new_count
            cached_metrics = new_metrics
        
        print(count, 'metrics updated')
        # 控制采集间隔,与Prometheus抓取间隔对齐(示例为10秒)
        sleep_duration = max(0, 10 - (time.time() - start_time))
        time.sleep(sleep_duration)

class TestExporter:
    def collect(self):
        with lock:
            # 直接返回缓存的指标
            for metric in cached_metrics:
                yield metric

if __name__ == '__main__':
    # 启动后台采集线程(守护线程随主进程退出)
    threading.Thread(target=background_collect, daemon=True).start()
    # 等待第一次采集完成,避免初始无指标
    while not cached_metrics:
        time.sleep(0.1)
    
    REGISTRY.register(TestExporter())
    start_http_server(8000)
    print("started server...")
    
    while True:
        time.sleep(1)

方案3:调整Prometheus超时配置(不推荐)

如果必须保留串行逻辑,可以修改Prometheus配置文件,延长抓取超时时间,但这会降低Prometheus整体抓取效率,仅作为临时方案:

scrape_configs:
  - job_name: 'test_exporter'
    scrape_interval: 10s
    scrape_timeout: 20s  # 设置为大于采集总耗时的值
    static_configs:
      - targets: ['localhost:8000']

内容的提问来源于stack exchange,提问作者Vamp thehacker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 16:45:42