Nagios中check_by_ssh模块报错:skip-stderr参数需为整数
I've run into this exact issue before—this error pops up because special characters in your command aren't properly escaped for Nagios' configuration file, causing it to misinterpret parts of your remote command as local arguments for check_by_ssh. Let's break down how to fix this:
Root Cause
The main problem is that characters like | (pipe) and single quotes in your remote command get parsed by the local shell when Nagios reads the config, instead of being passed through to the remote host. This confuses check_by_ssh into thinking you're passing invalid arguments (like the skip-stderr flag mentioned in the error). Also, note your original awk command is truncated (print ...), so we'll assume a complete status-output logic for the fix.
Step-by-Step Fixes
1. Properly Escape Special Characters in commands.cfg
In Nagios' commands.cfg, you need to escape pipes and handle single quotes correctly. Wrap your entire remote command in single quotes, and escape internal single quotes using '\'' (a shell trick to escape single quotes inside a single-quoted string).
Here's the corrected command definition:
define command{ command_name check_established_connections command_line /usr/local/nagios/libexec/check_by_ssh -l fuseadmin -H $HOSTADDRESS$ -C 'netstat -punta | grep -i ESTABLISHED | wc -l | awk '\''{if ($0>2500) {print "CRITICAL - Established connections: "$0"|connections="$0; exit 2} else if ($0>2000) {print "WARNING - Established connections: "$0"|connections="$0; exit 1} else {print "OK - Established connections: "$0"|connections="$0; exit 0}}'\'' }
Key changes:
- Replaced hardcoded
<hostname>with Nagios'$HOSTADDRESS$macro for dynamic host targeting - Wrapped the remote command in single quotes to isolate it from the local shell
- Escaped awk's internal single quotes with
'\''so they're passed to the remote host - Added complete exit codes (2 for critical, 1 for warning, 0 for OK) which Nagios relies on
2. Test the Command Manually First
Before updating the config, test the escaped command directly on your Nagios server to make sure it works:
/usr/local/nagios/libexec/check_by_ssh -l fuseadmin -H your_actual_hostname -C 'netstat -punta | grep -i ESTABLISHED | wc -l | awk '\''{if ($0>2500) {print "CRITICAL - Established connections: "$0"|connections="$0; exit 2} else if ($0>2000) {print "WARNING - Established connections: "$0"|connections="$0; exit 1} else {print "OK - Established connections: "$0"|connections="$0; exit 0}}'\''
If this returns the correct status and output, proceed to update commands.cfg and restart Nagios:
systemctl restart nagios
3. Better Alternative: Use a Remote Script
For easier maintenance (and fewer escape headaches), create a simple shell script on the remote host instead of cramming all logic into the check_by_ssh command.
- On the remote host, create
/usr/local/bin/check_established.shwith this content:
#!/bin/bash COUNT=$(netstat -punta | grep -i ESTABLISHED | wc -l) if [ "$COUNT" -gt 2500 ]; then echo "CRITICAL - Established connections: $COUNT|connections=$COUNT" exit 2 elif [ "$COUNT" -gt 2000 ]; then echo "WARNING - Established connections: $COUNT|connections=$COUNT" exit 1 else echo "OK - Established connections: $COUNT|connections=$COUNT" exit 0 fi
- Make it executable:
chmod +x /usr/local/bin/check_established.sh
- Update your Nagios command to call this script:
define command{ command_name check_established_connections command_line /usr/local/nagios/libexec/check_by_ssh -l fuseadmin -H $HOSTADDRESS$ -C '/usr/local/bin/check_established.sh' }
This approach is cleaner, easier to debug, and avoids all the escape hassle.
内容的提问来源于stack exchange,提问作者AJ_NOVICE

