Prometheus Pushgateway如何实现Nuclio无服务请求数正确累加
问题原因分析
首先明确你的第一个疑问:是的,每次调用新建Counter对象会将本地累加值重置为0。Counter、Gauge这类指标对象默认在当前进程内存中初始化,Nuclio函数每次被调用时Pod重建、进程上下文完全重置,内存中没有历史计数,你每次调用inc()都是在初始值0的基础上加1,所以本地的Counter值永远是1。
另外你遇到pushadd_to_gateway也未生效的核心原因是分组标签冲突/唯一实例标签导致新序列生成:
push_to_gateway默认会直接覆盖网关中相同job+相同分组标签的指标值,你每次推1自然就一直是1pushadd_to_gateway如果没有指定固定的分组key,默认会自动携带当前主机名作为instance标签,每次Pod重建instance值变化后就会生成多条独立的时间序列,每条值都是1,你直接查询my_requests的时候如果没做聚合,就只会看到值为1的序列。
正确实现方案
针对这类短生命周期/无状态临时任务的指标累加,推荐两种可行方案:
方案1:正确使用Pushgateway做累加
核心逻辑是:固定分组标签,每次调用时先从Pushgateway拉取当前已有的计数值,在旧值基础上累加后再推回网关,避免本地初始值为0的问题。
代码示例:
from prometheus_client import CollectorRegistry, Counter, pushadd_to_gateway import requests def count_request(): # 固定分组标签,不要用动态的instance,避免生成多序列 group_key = {"instance": "nuclio-request-counter"} gateway_addr = "localhost:8082" job_name = "countJob" # 第一步:先从Pushgateway拉取当前已有的计数值 current_val = 0 try: # 查询对应job和分组标签的指标值 resp = requests.get(f"http://{gateway_addr}/metrics") for line in resp.text.split("\n"): if line.startswith('my_requests{') and 'job="countJob"' in line and 'instance="nuclio-request-counter"' in line: current_val = int(float(line.split(" ")[-1])) break except Exception: # 第一次查询不存在的话默认0 pass # 第二步:新建registry和Counter,基于旧值累加 registry = CollectorRegistry() c = Counter('my_requests', 'HTTP Requests Count', ['method', 'endpoint'], registry=registry) # 此处可根据当前实际请求类型修改标签值 c.labels(method='get', endpoint='/').inc(current_val + 1) # 第三步:推送到网关,指定固定分组key pushadd_to_gateway( gateway_addr, job=job_name, grouping_key=group_key, registry=registry )
注意:如果需要统计不同标签的请求数,拉取指标的时候需要按标签维度分别匹配旧值再累加
方案2:使用外部存储做计数中转
如果请求量比较大,拉取再推送的方式有性能问题,可以用Redis这类轻量外部存储存计数,定期把计数推送到Prometheus,或者直接让Prometheus对接Redis Exporter拉取指标。
简化示例:
import redis from prometheus_client import CollectorRegistry, Counter, push_to_gateway # 初始化Redis连接,建议封装为全局初始化逻辑 r = redis.Redis(host="你的Redis服务地址", port=6379, db=0, decode_responses=True) def count_request(): # 每次请求直接给Redis对应的key加1,原子操作不会有并发问题 r.incr("my_requests:get:/") # 按不同标签维度统计可存为不同key # r.incr("my_requests:post:/submit") # 可选:定期推送到Pushgateway或者直接让Prometheus拉取Redis的指标 registry = CollectorRegistry() c = Counter('my_requests', 'HTTP Requests Count', ['method', 'endpoint'], registry=registry) current_val = int(r.get("my_requests:get:/") or 0) c.labels(method='get', endpoint='/').inc(current_val) push_to_gateway("localhost:8082", job="countJob", registry=registry)
额外注意事项
- 如果Nuclio函数支持常驻进程模式,优先开启常驻,就可以在进程内存中保留Counter对象,不需要每次重建,直接inc后推或者让Prometheus拉取指标即可
- 多实例并发调用的场景下,方案1的拉取再推送逻辑可能会有计数丢失,建议加分布式锁或者直接用方案2的Redis原子incr操作
内容的提问来源于stack exchange,提问作者Masterbuilder
相关产品推荐
相关产品推荐

