自定义Exporter处理时长超Prometheus抓取间隔时无指标返回
问题根源
Prometheus默认抓取超时为10秒,当你的collect()方法执行时间超过这个阈值时,Prometheus会主动终止请求,导致无法获取任何指标,因此图表呈现空白。你的代码中串行执行3次check_stream,每次耗时5秒,总时长15秒,超过了默认超时限制。
解决方案
方案1:并行执行采集任务,缩短总耗时
将串行的check_stream调用改为并行执行,利用线程池减少整体耗时,确保总时长低于Prometheus超时阈值。
修改后的代码:
#!/usr/bin/env python3 from prometheus_client import start_http_server, REGISTRY from prometheus_client.metrics_core import GaugeMetricFamily import time from concurrent.futures import ThreadPoolExecutor count = 0 # 初始化线程池,最多同时运行3个线程 executor = ThreadPoolExecutor(max_workers=3) class TestExporter: def collect(self): global count count += 5 # 并行提交3个采集任务 futures = [executor.submit(self.check_stream, count, i) for i in range(3)] # 依次获取任务结果并返回 for future in futures: metric = future.result() yield metric def check_stream(self, count, host): time.sleep(5) metric = GaugeMetricFamily('aa_stream_test', 'testing stream delete', labels=['stream']) if count > 100: metric.add_metric(['B', str(host)], count) else: metric.add_metric(['A', str(host)], count) print(count, 'registering . . .') return metric if __name__ == '__main__': REGISTRY.register(TestExporter()) start_http_server(8000) print("started server...") while True: time.sleep(1)
方案2:异步预采集+缓存,快速响应抓取请求
如果单任务耗时过长,并行仍无法满足超时要求,可以在后台线程定期采集指标并缓存,collect()方法直接返回缓存结果,确保响应时间极短。
修改后的代码:
#!/usr/bin/env python3 from prometheus_client import start_http_server, REGISTRY from prometheus_client.metrics_core import GaugeMetricFamily import time import threading count = 0 cached_metrics = [] # 线程锁保证缓存更新的线程安全 lock = threading.Lock() def background_collect(): global count, cached_metrics while True: start_time = time.time() new_count = count + 5 new_metrics = [] # 并行采集指标(也可根据需求串行) for i in range(3): time.sleep(5) metric = GaugeMetricFamily('aa_stream_test', 'testing stream delete', labels=['stream']) if new_count > 100: metric.add_metric(['B', str(i)], new_count) else: metric.add_metric(['A', str(i)], new_count) new_metrics.append(metric) # 加锁更新缓存 with lock: count = new_count cached_metrics = new_metrics print(count, 'metrics updated') # 控制采集间隔,与Prometheus抓取间隔对齐(示例为10秒) sleep_duration = max(0, 10 - (time.time() - start_time)) time.sleep(sleep_duration) class TestExporter: def collect(self): with lock: # 直接返回缓存的指标 for metric in cached_metrics: yield metric if __name__ == '__main__': # 启动后台采集线程(守护线程随主进程退出) threading.Thread(target=background_collect, daemon=True).start() # 等待第一次采集完成,避免初始无指标 while not cached_metrics: time.sleep(0.1) REGISTRY.register(TestExporter()) start_http_server(8000) print("started server...") while True: time.sleep(1)
方案3:调整Prometheus超时配置(不推荐)
如果必须保留串行逻辑,可以修改Prometheus配置文件,延长抓取超时时间,但这会降低Prometheus整体抓取效率,仅作为临时方案:
scrape_configs: - job_name: 'test_exporter' scrape_interval: 10s scrape_timeout: 20s # 设置为大于采集总耗时的值 static_configs: - targets: ['localhost:8000']
内容的提问来源于stack exchange,提问作者Vamp thehacker
相关产品推荐
相关产品推荐

