无需Prometheus或外部服务,如何基于Nginx指标及自定义指标(RPS、活跃连接)扩展应用?
Alright, let's break this down for you—no Prometheus, no external monitoring bloat, just straightforward ways to scale your app based on custom metrics (RPS/active connections) or Nginx data. Here are the practical, self-contained solutions:
1. Scaling with Your App's Custom API Metrics
Since your web app exposes its own metrics API, we can leverage that directly without any third-party tools.
Kubernetes Custom HPA + Local Metric Adapter
Kubernetes' Horizontal Pod Autoscaler (HPA) doesn't require Prometheus—you can build a tiny, local metric adapter to bridge your app's API to K8s' custom metrics API. Here's how:
- Write a simple service (Python/Go/whatever you prefer) that polls your app's metrics endpoint (e.g.,
/api/metrics) every 10-30 seconds, pulls RPS/active connection counts, and formats that data to match K8s' custom metrics schema. - Deploy this adapter as a pod in your cluster, and register it with K8s' custom metrics API.
- Configure your HPA to use this custom metric for scaling.
Example HPA config snippet:
apiVersion: autoscaling/v2 kind: HorizontalPodAutoscaler metadata: name: your-app-hpa spec: scaleTargetRef: apiVersion: apps/v1 kind: Deployment name: your-app-deployment minReplicas: 2 maxReplicas: 10 metrics: - type: Pods pods: metric: name: app_active_connections target: type: AverageValue averageValue: 500
This tells K8s to scale your deployment when the average active connections per pod hits 500.
Script-Driven Cloud Auto-Scaling (AWS/Azure/Alibaba Cloud)
If you're using a cloud provider's Auto Scaling Group (ASG), a simple scheduled script can do the trick. Use your cloud's SDK to tie your app's metrics directly to ASG scaling:
- Write a shell/Python script that calls your app's metrics API, grabs the current RPS/active connections, then checks against your thresholds.
- Use cloud functions (Lambda/Azure Functions/Alibaba Cloud Function Compute) to run this script on a schedule (e.g., every 15 seconds).
Example bash script for AWS:
#!/bin/bash # Pull active connections from your app ACTIVE_CONNS=$(curl -s http://your-app/api/metrics | jq '.active_connections') # Get current ASG desired capacity CURRENT_CAPACITY=$(aws autoscaling describe-auto-scaling-groups --auto-scaling-group-names your-asg --query 'AutoScalingGroups[0].DesiredCapacity' --output text) # Scale out if connections are high and we haven't hit max if [ $ACTIVE_CONNS -gt 800 ] && [ $CURRENT_CAPACITY -lt 10 ]; then aws autoscaling set-desired-capacity --auto-scaling-group-name your-asg --desired-capacity $((CURRENT_CAPACITY + 2)) # Scale in if connections are low and we're above min elif [ $ACTIVE_CONNS -lt 200 ] && [ $CURRENT_CAPACITY -gt 2 ]; then aws autoscaling set-desired-capacity --auto-scaling-group-name your-asg --desired-capacity $((CURRENT_CAPACITY - 1)) fi
Docker Swarm Custom Scaling Script
For Docker Swarm, you can write a script that talks directly to the Swarm API to adjust service replicas based on your app's metrics:
- Use the Docker SDK (for Python/Go) to fetch your service's current replica count.
- Poll your app's metrics API, compare against thresholds, and call the Swarm API to scale up/down.
Example Python snippet:
import requests import docker client = docker.from_env() # Fetch metrics from your app metrics = requests.get("http://your-app/api/metrics").json() active_conns = metrics["active_connections"] # Get current replica count service = client.services.get("your-app-service") current_replicas = service.attrs["Spec"]["Mode"]["Replicated"]["Replicas"] # Scale logic if active_conns > 600 and current_replicas < 8: service.scale(current_replicas + 2) elif active_conns < 150 and current_replicas > 2: service.scale(current_replicas - 1)
Run this script as a scheduled task in your Swarm cluster (e.g., a container that runs the script every 20 seconds).
2. Scaling Based on Nginx Metrics (No Prometheus)
Nginx has built-in ways to expose metrics—we can tap into those directly for scaling.
Nginx stub_status + K8s Custom HPA Adapter
Nginx's stub_status module gives you basic metrics (active connections, request counts) out of the box. Here's how to use it:
- Enable
stub_statusin your Nginx config:
server { listen 80; location /nginx_status { stub_status on; allow 127.0.0.1; # Restrict access to internal services deny all; } }
This returns output like:
Active connections: 300 server accepts handled requests 12345 12345 456789 Reading: 0 Writing: 10 Waiting: 290
- Build a simple adapter to parse this output (extract active connections or calculate RPS from request counts) and expose it as a K8s custom metric. Then configure your HPA to use this metric—same as the custom API metric setup above.
Nginx Access Log Parsing + Cloud Scaling Script
If you don't want to use stub_status, you can parse Nginx's access log in real-time to calculate RPS:
- Use a shell script with
tail,awk, andbcto compute RPS over a window (e.g., 10 seconds):
#!/bin/bash # Calculate RPS over the last 10 seconds RPS=$(tail -n 1000 /var/log/nginx/access.log | awk '{print $4}' | sed 's/\[//g' | awk -F: '{print $4":"$5}' | sort | uniq -c | awk '{sum+=$1} END {print sum/10}') # Trigger cloud scaling if RPS exceeds threshold if [ $(echo "$RPS > 100" | bc) -eq 1 ]; then aws autoscaling set-desired-capacity --auto-scaling-group-name your-asg --desired-capacity 8 fi
Run this script on a schedule, or use inotifywait to trigger it whenever the log file updates.
Nginx Plus Built-In Auto-Scaling
If you're using Nginx Plus (the commercial version), it has native auto-scaling integration with Kubernetes, AWS ASG, and other platforms. You can configure it to scale your backend services based on its own collected metrics (active connections, RPS) without any external tools:
upstream app { zone app 64k; server app-1:80; server app-2:80; auto_scale on; scale_target "your-k8s-deployment" namespace "default"; scale_thresholds { active_connections 500; rps 200; } }
Nginx Plus will directly talk to the K8s API to adjust your deployment's replica count when thresholds are hit.
Quick Tips to Avoid Headaches
- Set cooldown periods: Prevent thrashing (constant scaling up/down) by adding cooldown rules (e.g., K8s HPA's
behaviorconfig, cloud ASG cooldown timers). - Handle failures: Add error handling to your scripts/adapters—if your app's metrics API goes down, don't trigger accidental scaling.
- Tune polling frequency: Polling every 10-30 seconds is a sweet spot—faster than that adds unnecessary load, slower means scaling reacts too late.
内容的提问来源于stack exchange,提问作者Zàf Mohammed

