You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS T3.small Linux实例CPU使用率达80%时无响应的问题排查求助

AWS T3.small Linux实例CPU使用率达80%时无响应的问题排查求助

Hey there, sorry to hear you're stuck with this frustrating issue—let's walk through some practical checks and fixes to get to the bottom of it:

  • Dig deeper into CPU usage details
    T3.small has 2 vCPUs, but a single-threaded process can max out one core even if total CPU sits at 80%. That's often enough to grind the system to a halt because critical services (like SSH) get starved for CPU time. Next time the campaign runs, try these:

    • Use htop (install it if needed) to monitor individual core usage instead of just total CPU. Look for any core hitting 100% usage.
    • Run pidstat -u 1 to track which specific process is eating up CPU—focus on your Google campaign-related tasks. Check if it's a single-threaded process hogging one core.
    • Set up CloudWatch alarms for per-vCPU usage (not just aggregate) to capture this behavior even when you can't SSH in.
  • Check for hidden memory/disk bottlenecks
    Sometimes CPU spikes are a symptom, not the cause. Even if total CPU is at 80%, your system might be choking on memory or disk IO:

    • Run free -h to check if you're running low on RAM. If memory is nearly full, the system might be thrashing swap space, which makes everything unresponsive (even if CPU doesn't hit 100%).
    • Use iostat -x 1 to monitor disk IO. Look for high %util (over 90%) or long await times—this means disk operations are backing up and blocking system tasks.
    • Check dmesg for OOM (Out of Memory) killer logs—sometimes the system kills processes quietly without showing up in general logs, which could disrupt SSH or other services.
  • Secure a fallback SSH access
    When the system is unresponsive, you can't diagnose it in real-time. Fix this by:

    • Keeping a persistent SSH session open with screen or tmux before the campaign runs. If new SSH connections fail, you can switch back to this session to check what's going on.
    • Tweak your SSH daemon config (/etc/ssh/sshd_config) to add ClientAliveInterval 30 and ClientAliveCountMax 3—this helps keep existing connections alive even under high load. Also check /var/log/secure or /var/log/auth.log for SSH-specific errors that might not show up in general logs.
  • Adjust process priorities
    If your Google campaign process is the culprit, lower its priority so critical services get CPU first:

    • When starting the campaign, launch it with nice -n 10 <your-campaign-command> to give it a lower priority.
    • If it's already running, use renice 10 <pid> to adjust its priority on the fly. This won't stop it from using CPU, but it ensures SSH and other system services aren't starved.
  • Test with temporary instance scaling
    Since this only happens every 15 days, try temporarily upgrading to a T3.medium (2 vCPUs, 4GB RAM) before the next campaign runs. If the issue disappears, it means the T3.small's resources (even with unlimited CPU credits) aren't sufficient to handle the campaign's peak load—maybe it's using more memory than you realize, or the single-core performance is the bottleneck.

Hope these tips help you track down the root cause! Let us know if you find any interesting clues from these checks.

备注:内容来源于stack exchange,提问作者hasnain hakim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.22 14:50:30