如何将catalina请求计数指标转为请求率用于HPA弹性扩缩容?
解决Tomcat累计请求数指标用于HPA扩缩容的问题
1. 将累计请求数转为请求率(推荐方案)
由于catalina_globalrequestprocessor_requestcount是累计型计数器(counter),直接用绝对值做阈值会导致持续触发扩容。正确的做法是计算其每秒请求率,用速率作为HPA的判断指标。
在GKE结合Managed Prometheus和Stackdriver Adapter的架构下,有两种实现方式:
方式一:通过Prometheus记录规则生成速率指标
在Managed Prometheus中创建记录规则,预先计算请求率并生成新的Gauge指标:
# prometheus-recording-rule.yaml apiVersion: monitoring.googleapis.com/v1 kind: PrometheusRule metadata: name: tomcat-request-rate namespace: my-ns spec: groups: - name: tomcat.rules rules: - record: tomcat_requests_per_second expr: rate(catalina_globalrequestprocessor_requestcount[1m]) labels: app: my-deployment
创建后,Stackdriver Adapter会自动同步这个新指标,之后修改HPA配置使用该指标:
apiVersion: autoscaling/v2 kind: HorizontalPodAutoscaler metadata: name: my-autoscale namespace: my-ns spec: maxReplicas: 3 metrics: - pods: metric: name: prometheus.googleapis.com|tomcat_requests_per_second|gauge target: averageValue: 5 # 每个Pod平均每秒5个请求,按需调整 type: AverageValue type: Pods minReplicas: 1 scaleTargetRef: apiVersion: apps/v1 kind: Deployment name: my-deployment
方式二:直接在HPA中使用PromQL表达式
无需预先创建规则,直接在HPA的External指标中指定速率计算的PromQL:
apiVersion: autoscaling/v2 kind: HorizontalPodAutoscaler metadata: name: my-autoscale namespace: my-ns spec: maxReplicas: 3 metrics: - external: metric: name: prometheus.googleapis.com|query|result selector: matchLabels: query: 'rate(catalina_globalrequestprocessor_requestcount{pod=~"my-deployment-.*"}[1m])' target: averageValue: "5" # 每秒请求数阈值 type: AverageValue type: External minReplicas: 1 scaleTargetRef: apiVersion: apps/v1 kind: Deployment name: my-deployment
注意:
[1m]是速率计算的时间窗口,可根据业务流量波动调整(比如[30s]适合快速波动的流量)。
2. 替代的Tomcat MBean指标
如果不想计算速率,可选用Tomcat暴露的其他Gauge类型指标作为扩缩容依据:
catalina_threadpool_currentthreadsbusiness:当前繁忙的线程数,直接反映Pod的处理负载,适合作为核心扩缩容指标(比如平均每个Pod繁忙线程数超过15时扩容)catalina_globalrequestprocessor_currentconnections:当前活跃连接数,体现实时流量压力catalina_memory_heap_usedpercent:堆内存使用率,你已验证过该指标的有效性,可结合请求率做复合指标扩缩容
内容的提问来源于stack exchange,提问作者brandizzi
相关产品推荐
相关产品推荐

