AWS T3.small Linux实例CPU使用率达80%时无响应的问题排查求助
Hey there, sorry to hear you're stuck with this frustrating issue—let's walk through some practical checks and fixes to get to the bottom of it:
Dig deeper into CPU usage details
T3.small has 2 vCPUs, but a single-threaded process can max out one core even if total CPU sits at 80%. That's often enough to grind the system to a halt because critical services (like SSH) get starved for CPU time. Next time the campaign runs, try these:- Use
htop(install it if needed) to monitor individual core usage instead of just total CPU. Look for any core hitting 100% usage. - Run
pidstat -u 1to track which specific process is eating up CPU—focus on your Google campaign-related tasks. Check if it's a single-threaded process hogging one core. - Set up CloudWatch alarms for per-vCPU usage (not just aggregate) to capture this behavior even when you can't SSH in.
- Use
Check for hidden memory/disk bottlenecks
Sometimes CPU spikes are a symptom, not the cause. Even if total CPU is at 80%, your system might be choking on memory or disk IO:- Run
free -hto check if you're running low on RAM. If memory is nearly full, the system might be thrashing swap space, which makes everything unresponsive (even if CPU doesn't hit 100%). - Use
iostat -x 1to monitor disk IO. Look for high%util(over 90%) or longawaittimes—this means disk operations are backing up and blocking system tasks. - Check
dmesgfor OOM (Out of Memory) killer logs—sometimes the system kills processes quietly without showing up in general logs, which could disrupt SSH or other services.
- Run
Secure a fallback SSH access
When the system is unresponsive, you can't diagnose it in real-time. Fix this by:- Keeping a persistent SSH session open with
screenortmuxbefore the campaign runs. If new SSH connections fail, you can switch back to this session to check what's going on. - Tweak your SSH daemon config (
/etc/ssh/sshd_config) to addClientAliveInterval 30andClientAliveCountMax 3—this helps keep existing connections alive even under high load. Also check/var/log/secureor/var/log/auth.logfor SSH-specific errors that might not show up in general logs.
- Keeping a persistent SSH session open with
Adjust process priorities
If your Google campaign process is the culprit, lower its priority so critical services get CPU first:- When starting the campaign, launch it with
nice -n 10 <your-campaign-command>to give it a lower priority. - If it's already running, use
renice 10 <pid>to adjust its priority on the fly. This won't stop it from using CPU, but it ensures SSH and other system services aren't starved.
- When starting the campaign, launch it with
Test with temporary instance scaling
Since this only happens every 15 days, try temporarily upgrading to a T3.medium (2 vCPUs, 4GB RAM) before the next campaign runs. If the issue disappears, it means the T3.small's resources (even with unlimited CPU credits) aren't sufficient to handle the campaign's peak load—maybe it's using more memory than you realize, or the single-core performance is the bottleneck.
Hope these tips help you track down the root cause! Let us know if you find any interesting clues from these checks.
备注:内容来源于stack exchange,提问作者hasnain hakim

