本地Nagios服务器通过NRPE监控AWS Linux实例遇连接超时求助
Hey Maniraj, let's work through this NRPE socket timeout issue step by step—this is super common when connecting to AWS instances, so we'll cover all the usual culprits first:
This is the most frequent root cause for this timeout:
- On your AWS instance's security group, add an inbound rule allowing TCP port 5666 (NRPE's default port) from your Nagios server's public/private IP. Avoid using
0.0.0.0/0for security—lock it to your Nagios server's specific IP. - Double-check your Nagios server's outbound rules (if applicable) to ensure it can initiate TCP connections to port 5666. Most default security groups allow all outbound traffic, but strict setups might block this.
Even if security groups are open, the instance's local firewall could be blocking traffic:
- For
ufwusers: Runsudo ufw statusto confirm there's a rule allowing port 5666 from your Nagios IP. If not, add it with:sudo ufw allow from [YOUR_NAGIOS_SERVER_IP] to any port 5666 - For
iptablesusers: Runsudo iptables -L -nto look for a rule likeACCEPT tcp -- [YOUR_NAGIOS_IP] 0.0.0.0/0 tcp dpt:5666. If missing, add it and save the rules:sudo iptables -A INPUT -s [YOUR_NAGIOS_IP] -p tcp --dport 5666 -j ACCEPT sudo iptables-save > /etc/iptables/rules.v4
Incorrect NRPE settings can block valid connections:
- Open the NRPE config file (usually
/usr/local/nagios/etc/nrpe.cfg) and verify:allowed_hosts: Must include your Nagios server's IP, e.g.,allowed_hosts=127.0.0.1,[YOUR_NAGIOS_IP](no spaces between entries)server_address: Either leave blank (to listen on all interfaces) or set it to the instance's private/public IPserver_port: Confirm it's set to 5666 (matches what you're using incheck_nrpe)
- Restart the NRPE service to apply changes:
sudo systemctl restart nrpe # Or for init.d systems: sudo /etc/init.d/nrpe restart
Skip NRPE entirely to rule out network issues:
- From your Nagios server, run
telnet [AWS_INSTANCE_IP] 5666. If connected, you'll see a response likeNRPE v2.15. If this times out, go back to security group/firewall checks. - Alternatively, use
nc:
A successful connection will shownc -zv [AWS_INSTANCE_IP] 5666succeeded!.
If network checks pass but NRPE still times out, make sure the service is active:
- On the AWS instance, run:
sudo systemctl status nrpe # Or check running processes: ps aux | grep nrpe - If it's not running, start it and enable auto-start:
sudo systemctl start nrpe sudo systemctl enable nrpe
If you're using private IPs to connect (same VPC or peered VPCs):
- Verify your VPC route tables allow traffic between the Nagios server and AWS instance
- Ensure VPC Network ACLs aren't blocking inbound/outbound TCP 5666 traffic
Start with the network checks first—9 times out of 10, it's a security group or firewall rule blocking the port. Let me know if any of these steps resolve your issue, or if you hit something unexpected!
内容的提问来源于stack exchange,提问作者Maniraj

