AWS实例卡顿无法SSH连接、网站失效,求非重启解决方案
Alright, let's figure out how to resolve your EC2 instance freeze—where SSH connections fail and your hosted site is down—without hitting the restart button in the AWS Console. I’ve dealt with similar issues before, so here are actionable, hands-on solutions to try:
1. Use AWS Systems Manager (SSM) Run Command
If your instance has the SSM Agent installed (most official AWS AMIs come with it pre-configured), you can send shell commands remotely to diagnose and fix issues:
- First, check which processes are hogging resources:
aws ssm send-command --instance-ids "your-instance-id" --document-name "AWS-RunShellScript" --parameters commands="top -b -n 1 | head -20; ps aux --sort=-%cpu | head -10" - Once you identify the unresponsive process, terminate it with:
aws ssm send-command --instance-ids "your-instance-id" --document-name "AWS-RunShellScript" --parameters commands="sudo kill -9 [PID-of-offending-process]" - If the instance still won’t recover, send a soft restart command (gentler than a console reboot):
aws ssm send-command --instance-ids "your-instance-id" --document-name "AWS-RunShellScript" --parameters commands="sudo shutdown -r now"
2. Connect via EC2 Serial Console
If you’ve enabled the EC2 Serial Console for your account and instance (supported in most regions), this lets you bypass SSH and connect directly to the instance’s serial port—perfect if network issues are blocking SSH:
- Start a serial session via CLI:
aws ec2 start-session --target "your-instance-id" - Once connected, log in with your instance’s OS credentials (you may need to set a password first for Linux instances if you only used SSH keys before).
- From here, you can check system logs, kill stuck processes, or initiate a soft reboot—all without touching the Console restart button.
3. Diagnose Network Blockages
Sometimes the issue isn’t the instance itself, but network rules blocking traffic:
- Verify your instance’s network interface is healthy:
Ensure theaws ec2 describe-network-interfaces --filters "Name=attachment.instance-id,Values=your-instance-id"Statusisin-useandAttachment.Statusisattached. - Check your security group allows SSH (port 22) and your site’s ports (80/443) for inbound traffic:
aws ec2 describe-security-groups --group-ids "your-security-group-id" - Confirm Network ACLs (NACLs) aren’t blocking necessary traffic (default NACLs allow all, but custom rules might cause issues):
aws ec2 describe-network-acls --filters "Name=association.subnet-id,Values=your-subnet-id"
4. Analyze CloudWatch Metrics & Logs
Before taking action, pinpoint the root cause using CloudWatch:
- Check CPU, memory, and disk I/O metrics to see if resources are maxed out:
# Check CPU utilization over the last hour aws cloudwatch get-metric-statistics --namespace AWS/EC2 --metric-name CPUUtilization --dimensions Name=InstanceId,Value=your-instance-id --start-time $(date -d "-1 hour" +%Y-%m-%dT%H:%M:%SZ) --end-time $(date +%Y-%m-%dT%H:%M:%SZ) --period 60 --statistics Average - Pull the instance’s console output to look for system errors, disk full messages, or process crashes:
aws ec2 get-console-output --instance-id "your-instance-id"
内容的提问来源于stack exchange,提问作者Saifullah khan

