You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无需Prometheus或外部服务,如何基于Nginx指标及自定义指标(RPS、活跃连接)扩展应用?

Alright, let's break this down for you—no Prometheus, no external monitoring bloat, just straightforward ways to scale your app based on custom metrics (RPS/active connections) or Nginx data. Here are the practical, self-contained solutions:

1. Scaling with Your App's Custom API Metrics

Since your web app exposes its own metrics API, we can leverage that directly without any third-party tools.

Kubernetes Custom HPA + Local Metric Adapter

Kubernetes' Horizontal Pod Autoscaler (HPA) doesn't require Prometheus—you can build a tiny, local metric adapter to bridge your app's API to K8s' custom metrics API. Here's how:

  • Write a simple service (Python/Go/whatever you prefer) that polls your app's metrics endpoint (e.g., /api/metrics) every 10-30 seconds, pulls RPS/active connection counts, and formats that data to match K8s' custom metrics schema.
  • Deploy this adapter as a pod in your cluster, and register it with K8s' custom metrics API.
  • Configure your HPA to use this custom metric for scaling.

Example HPA config snippet:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: your-app-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: your-app-deployment
  minReplicas: 2
  maxReplicas: 10
  metrics:
  - type: Pods
    pods:
      metric:
        name: app_active_connections
      target:
        type: AverageValue
        averageValue: 500

This tells K8s to scale your deployment when the average active connections per pod hits 500.

Script-Driven Cloud Auto-Scaling (AWS/Azure/Alibaba Cloud)

If you're using a cloud provider's Auto Scaling Group (ASG), a simple scheduled script can do the trick. Use your cloud's SDK to tie your app's metrics directly to ASG scaling:

  • Write a shell/Python script that calls your app's metrics API, grabs the current RPS/active connections, then checks against your thresholds.
  • Use cloud functions (Lambda/Azure Functions/Alibaba Cloud Function Compute) to run this script on a schedule (e.g., every 15 seconds).

Example bash script for AWS:

#!/bin/bash
# Pull active connections from your app
ACTIVE_CONNS=$(curl -s http://your-app/api/metrics | jq '.active_connections')
# Get current ASG desired capacity
CURRENT_CAPACITY=$(aws autoscaling describe-auto-scaling-groups --auto-scaling-group-names your-asg --query 'AutoScalingGroups[0].DesiredCapacity' --output text)

# Scale out if connections are high and we haven't hit max
if [ $ACTIVE_CONNS -gt 800 ] && [ $CURRENT_CAPACITY -lt 10 ]; then
  aws autoscaling set-desired-capacity --auto-scaling-group-name your-asg --desired-capacity $((CURRENT_CAPACITY + 2))
# Scale in if connections are low and we're above min
elif [ $ACTIVE_CONNS -lt 200 ] && [ $CURRENT_CAPACITY -gt 2 ]; then
  aws autoscaling set-desired-capacity --auto-scaling-group-name your-asg --desired-capacity $((CURRENT_CAPACITY - 1))
fi

Docker Swarm Custom Scaling Script

For Docker Swarm, you can write a script that talks directly to the Swarm API to adjust service replicas based on your app's metrics:

  • Use the Docker SDK (for Python/Go) to fetch your service's current replica count.
  • Poll your app's metrics API, compare against thresholds, and call the Swarm API to scale up/down.

Example Python snippet:

import requests
import docker

client = docker.from_env()
# Fetch metrics from your app
metrics = requests.get("http://your-app/api/metrics").json()
active_conns = metrics["active_connections"]
# Get current replica count
service = client.services.get("your-app-service")
current_replicas = service.attrs["Spec"]["Mode"]["Replicated"]["Replicas"]

# Scale logic
if active_conns > 600 and current_replicas < 8:
    service.scale(current_replicas + 2)
elif active_conns < 150 and current_replicas > 2:
    service.scale(current_replicas - 1)

Run this script as a scheduled task in your Swarm cluster (e.g., a container that runs the script every 20 seconds).


2. Scaling Based on Nginx Metrics (No Prometheus)

Nginx has built-in ways to expose metrics—we can tap into those directly for scaling.

Nginx stub_status + K8s Custom HPA Adapter

Nginx's stub_status module gives you basic metrics (active connections, request counts) out of the box. Here's how to use it:

  • Enable stub_status in your Nginx config:
server {
    listen 80;
    location /nginx_status {
        stub_status on;
        allow 127.0.0.1; # Restrict access to internal services
        deny all;
    }
}

This returns output like:

Active connections: 300 
server accepts handled requests
  12345 12345 456789
Reading: 0 Writing: 10 Waiting: 290 
  • Build a simple adapter to parse this output (extract active connections or calculate RPS from request counts) and expose it as a K8s custom metric. Then configure your HPA to use this metric—same as the custom API metric setup above.

Nginx Access Log Parsing + Cloud Scaling Script

If you don't want to use stub_status, you can parse Nginx's access log in real-time to calculate RPS:

  • Use a shell script with tail, awk, and bc to compute RPS over a window (e.g., 10 seconds):
#!/bin/bash
# Calculate RPS over the last 10 seconds
RPS=$(tail -n 1000 /var/log/nginx/access.log | awk '{print $4}' | sed 's/\[//g' | awk -F: '{print $4":"$5}' | sort | uniq -c | awk '{sum+=$1} END {print sum/10}')

# Trigger cloud scaling if RPS exceeds threshold
if [ $(echo "$RPS > 100" | bc) -eq 1 ]; then
    aws autoscaling set-desired-capacity --auto-scaling-group-name your-asg --desired-capacity 8
fi

Run this script on a schedule, or use inotifywait to trigger it whenever the log file updates.

Nginx Plus Built-In Auto-Scaling

If you're using Nginx Plus (the commercial version), it has native auto-scaling integration with Kubernetes, AWS ASG, and other platforms. You can configure it to scale your backend services based on its own collected metrics (active connections, RPS) without any external tools:

upstream app {
    zone app 64k;
    server app-1:80;
    server app-2:80;
    auto_scale on;
    scale_target "your-k8s-deployment" namespace "default";
    scale_thresholds {
        active_connections 500;
        rps 200;
    }
}

Nginx Plus will directly talk to the K8s API to adjust your deployment's replica count when thresholds are hit.


Quick Tips to Avoid Headaches

  • Set cooldown periods: Prevent thrashing (constant scaling up/down) by adding cooldown rules (e.g., K8s HPA's behavior config, cloud ASG cooldown timers).
  • Handle failures: Add error handling to your scripts/adapters—if your app's metrics API goes down, don't trigger accidental scaling.
  • Tune polling frequency: Polling every 10-30 seconds is a sweet spot—faster than that adds unnecessary load, slower means scaling reacts too late.

内容的提问来源于stack exchange,提问作者Zàf Mohammed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 11:34:14