You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何阻止K8S HPA在负载降低时删除正在执行任务的Pod?

Great question—this is a super common pain point when pairing Sidekiq with Kubernetes HPA, since Sidekiq doesn’t natively signal to Kubernetes when a pod is mid-job. Let’s walk through exactly how to fix this so you only scale down when queues are empty and no pods are actively processing work:

解决方案步骤

1. 扩展Prometheus Adapter暴露Sidekiq运行中任务指标

First, you need to make Kubernetes aware of which Sidekiq pods are actively running jobs. Assuming you’re already using a Sidekiq exporter (like prometheus-sidekiq-exporter), it should output a metric like sidekiq_process_busy (tracking in-progress jobs per pod).

Add a rule to your Prometheus Adapter config to expose this as a Kubernetes-native metric:

rules:
- seriesQuery: 'sidekiq_process_busy{namespace!="",pod!=""}'
  resources:
    overrides:
      namespace: {resource: "namespace"}
      pod: {resource: "pod"}
  name:
    matches: "^sidekiq_process_busy$"
    as: "sidekiq_running_jobs"
  metricsQuery: 'sum(sidekiq_process_busy{pod!="",namespace!=""}) by (pod, namespace)'

This lets Kubernetes fetch a sidekiq_running_jobs metric for each pod—where a value of 0 means the pod has no active tasks.

2. Configure HPA to Only Scale Down When Safe

Default HPA logic only looks at queue length, but we need to combine that with pod activity to control scaling. Here are two actionable approaches:

Approach A: Composite Metric for "Safe to Scale Down"

Create a custom Prometheus metric that only returns 0 when both the total queue size is 0 and all pods have no running jobs. Expose this via Prometheus Adapter as sidekiq_can_scale_down.

The PromQL for this metric would look like:

max(sidekiq_queue_size{queue="your-target-queue"} or vector(0)) + max(sidekiq_running_jobs{namespace="your-namespace"} or vector(0))

Then update your HPA to use this metric to gate scaling down, alongside your existing expansion rule:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: sidekiq-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: sidekiq-deployment
  minReplicas: 1
  maxReplicas: 10
  metrics:
  # Expansion logic: Scale up when average jobs per pod exceeds 100
  - type: Pods
    pods:
      metric:
        name: sidekiq_queue_size
      target:
        type: AverageValue
        averageValue: 100
  # Scale-down gate: Only allow scaling when sidekiq_can_scale_down equals 0
  - type: External
    external:
      metric:
        name: sidekiq_can_scale_down
      target:
        type: Value
        value: 0

HPA will only consider scaling down when both metrics are satisfied—queue empty, no active jobs.

Approach B: HPA Behavior + Pod Activity Metrics

Use HPA v2’s behavior field to slow scaling down, and add a pod-level metric to ensure we only scale down when pods are idle:

spec:
  behavior:
    scaleDown:
      stabilizationWindowSeconds: 300 # Wait 5 mins before scaling down to let jobs finish
      policies:
      - type: Pods
        value: 1
        periodSeconds: 60 # Only delete 1 pod per minute
      selectPolicy: Min
  metrics:
  # Expansion logic (unchanged)
  - type: Pods
    pods:
      metric:
        name: sidekiq_queue_size
      target:
        type: AverageValue
        averageValue: 100
  # Scale-down check: Only proceed if average running jobs per pod is 0
  - type: Pods
    pods:
      metric:
        name: sidekiq_running_jobs
      target:
        type: AverageValue
        averageValue: 0

This ensures HPA won’t scale down until all pods are idle, and gives existing jobs time to complete before any pods are terminated.

3. Enforce Sidekiq Graceful Shutdown

Even with HPA safeguards, you need to make sure pods don’t drop mid-job if they do get terminated. Add these settings to your Sidekiq deployment:

  • A long terminationGracePeriodSeconds (e.g., 300 seconds) to give Sidekiq time to finish tasks.
  • A preStop hook to tell Sidekiq to stop accepting new jobs and wrap up active ones:
lifecycle:
  preStop:
    exec:
      command: ["bundle", "exec", "sidekiqctl", "quiet", "/var/run/sidekiq.pid"]

4. Optional: Pod Disruption Budget (PDB)

Add a PDB to prevent accidental mass termination of pods, even if HPA misfires:

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: sidekiq-pdb
spec:
  minAvailable: 1
  selector:
    matchLabels:
      app: sidekiq

This guarantees at least one Sidekiq pod stays running at all times.

Final Notes

The key here is tying Kubernetes scaling decisions directly to Sidekiq’s actual workload state—empty queues and idle pods. Combine the metric exposure, HPA logic, and graceful shutdown settings, and you’ll eliminate terminated in-progress jobs entirely.

内容的提问来源于stack exchange,提问作者Hb_1993

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:09:22