AWS新手求助:如何为PHP-Video-Transcoder配置Auto Scaling
Hey there! Let's break down how to set up Auto Scaling for your PHP Video Transcoder app step by step—since you're already using RabbitMQ for queuing, that's a perfect foundation for scaling based on actual task load. Here's what you need to do:
Before setting up Auto Scaling, you need to fix any single-instance bottlenecks so new instances can seamlessly join your workflow:
- Move RabbitMQ to a managed or clustered service: Your current single-instance RabbitMQ will become a bottleneck if you scale out. Use AWS MQ (managed RabbitMQ) or set up a RabbitMQ cluster so all consumer instances can connect to the queue reliably.
- Switch to shared storage for media: Stop storing uploaded/transcoded videos on instance local disk. Use Amazon S3 instead—configure your PHP upload service to save videos directly to S3, and update your FFMPEG workers to pull source files from S3 and push finished MP4s back to S3.
- Standardize worker configuration: Package your Supervisor setup, PHP dependencies, FFMPEG, and worker code into an Amazon Machine Image (AMI). Alternatively, use user data scripts to automatically install/configure these components when a new instance launches. For example, a user data snippet might:
# Install dependencies apt-get update && apt-get install -y php ffmpeg supervisor rabbitmq-client # Clone your worker code git clone https://github.com/riyaskp/PHP-Video-Transcoder /opt/transcoder # Copy Supervisor config (8 workers for 8-core instances) cp /opt/transcoder/supervisor/transcode-worker.conf /etc/supervisor/conf.d/ # Start Supervisor supervisorctl reread && supervisorctl update && supervisorctl start transcode-worker:*
This is the core component that manages your worker instances:
- Build a launch template/configuration: Use your prepped AMI or user data script to define what each new instance looks like. Make sure to attach an IAM role with permissions for:
- Accessing S3
- Connecting to AWS MQ/RabbitMQ
- Sending metrics to CloudWatch
- Interacting with Auto Scaling lifecycle hooks
- Set ASG bounds: Define a minimum instance count (e.g., 1, to ensure at least one worker is always running), maximum instance count (e.g., 10, based on your budget and peak load), and initial desired count (1).
- Network setup: Place the ASG in private subnets within your VPC, and configure security groups to allow outbound traffic to S3, RabbitMQ (port 5672), and CloudWatch.
Your scaling should trigger based on two key signals—RabbitMQ queue depth (most critical, since it represents pending transcode tasks) and instance CPU usage (as a backup):
3.1 Send RabbitMQ Queue Depth to CloudWatch
CloudWatch doesn't natively track RabbitMQ metrics, so you'll need a small script to push queue depth as a custom metric. Add this script to your worker instances (via AMI or user data) and run it every minute with cron:
#!/bin/bash QUEUE_NAME="transcode-jobs" # Replace with your actual queue name REGION="us-east-1" # Replace with your AWS region # Get queue depth from RabbitMQ MESSAGE_COUNT=$(rabbitmqctl list_queues name messages | grep "$QUEUE_NAME" | awk '{print $2}') # Push to CloudWatch aws cloudwatch put-metric-data \ --namespace "RabbitMQ" \ --metric-name "QueueDepth" \ --dimensions QueueName="$QUEUE_NAME" \ --value "$MESSAGE_COUNT" \ --region "$REGION"
3.2 Create Scaling Policies
- Scale-out policy: When the
RabbitMQ/QueueDepthmetric exceeds a threshold (e.g., 100 pending tasks) for 5 minutes, add 1 instance (or scale proportionally, like +1 instance per 50 extra tasks). You can also add a secondary trigger if CPU usage stays above 70% for 5 minutes. - Scale-in policy: When queue depth drops below a threshold (e.g., 10 tasks) for 10 minutes, remove 1 instance, down to your minimum count.
When Auto Scaling terminates an instance, you don't want to drop in-progress transcode tasks. Set up a lifecycle hook in your ASG to trigger a shutdown script:
#!/bin/bash INSTANCE_ID=$(curl -s http://169.254.169.254/latest/meta-data/instance-id) ASG_NAME="your-transcode-asg" HOOK_NAME="terminate-worker-hook" # Stop Supervisor from accepting new tasks supervisorctl stop transcode-worker:* # Wait for all running workers to finish while supervisorctl status | grep -q RUNNING; do sleep 10 done # Notify Auto Scaling it's safe to terminate the instance aws autoscaling complete-lifecycle-action \ --lifecycle-hook-name "$HOOK_NAME" \ --auto-scaling-group-name "$ASG_NAME" \ --lifecycle-action-result CONTINUE \ --instance-id "$INSTANCE_ID"
- Simulate peak load: Upload a batch of videos to fill your RabbitMQ queue beyond your scale-out threshold, and verify that the ASG launches new instances automatically.
- Simulate low load: Wait for all tasks to complete, then confirm the ASG scales back down to your minimum instance count.
- Test graceful shutdown: Manually terminate a worker instance and check that in-progress tasks finish before the instance is terminated.
内容的提问来源于stack exchange,提问作者Riyas Kp

