基于inotify的水下无人机日志监控bash脚本故障与优化咨询
Hey there, let's tackle your log monitoring issues one by one—starting with the core problems in your original script, then the SSH-related exit bug, and wrapping up with solid optimizations.
1. Fixing Missing Log Lines (Inotify & Concurrent File Access)
Your script misses lines likely due to two key issues:
- Gaps between
inotifywaitcalls: Each time you invokeinotifywaitin a loop, there's a small window where log lines could be written without being detected. - Unreliable
tailusage: Usingtail -n2 | head -n1to grab the last line is error-prone, especially if writes are partial or concurrent processes are accessing the file.
Solution:
Switch to a continuous file-tracking approach that combines tail -F (follows files even if rotated) with inotify for lifecycle events:
monitor_log() { local logfile=$1 # Track new lines in real-time, no gaps tail -n0 -F "$logfile" | while read -r line; do # Check if line matches your target pattern if [[ "$line" =~ "YOUR_TRIGGER_PATTERN" ]]; then # Execute your on-modify action here echo "Triggered action for line: $line" fi done & local tail_pid=$! # Wait for close event to clean up the tail process inotifywait -q -e close_write "$logfile" kill $tail_pid } # Main loop: spawn a background monitor for each new log file while read -r new_log; do monitor_log "log_directory/$new_log" & done < <(inotifywait -q -m -e create --format '%f' log_directory)
This way, tail -F actively tracks all new writes regardless of other processes accessing the file, eliminating missed lines.
2. Fixing Timeout Capture Logic
Your original script had a variable name typo (exidCode instead of exitCode) which broke timeout detection. Even with that fixed, using inotifywait -t in a loop is unreliable for long-running timeouts.
Solution:
Track timestamps manually to detect 6-hour inactivity:
monitor_log_with_timeout() { local logfile=$1 local timeout=21600 # 6 hours in seconds local last_activity=$(date +%s) # Track new lines tail -n0 -F "$logfile" | while read -r line; do last_activity=$(date +%s) # Process line as needed if [[ "$line" =~ "RESUMING MISSION" ]]; then # Handle mission resume action fi done & local tail_pid=$! # Check for timeout or close event while true; do local now=$(date +%s) if (( now - last_activity > timeout )); then echo "Timeout reached for $logfile" kill $tail_pid break fi # Poll for close event with short timeout to check inactivity if inotifywait -q -t 60 -e close_write "$logfile"; then kill $tail_pid break fi done }
This combines active line tracking with periodic timeout checks, ensuring you catch both close events and inactivity timeouts.
3. Monitoring Multiple Drones Simultaneously
Your original script ran serially—each log file blocked the main loop. The fix is to run each log monitor in the background (as shown in the examples above).
PHP vs Bash for Multi-Drone Monitoring:
- Bash: Works well for small numbers of drones, but managing dozens of background processes can get messy (you'll need to track PIDs and clean up stale processes).
- PHP with PECL inotify: A better scalable solution. PHP's inotify extension lets you add multiple watchers to a single event loop, so you can monitor all log files in one process without spawning tons of background jobs. Example snippet:
$logDir = '/path/to/log_directory'; $inotify = inotify_init(); stream_set_blocking($inotify, false); // Watch directory for new files inotify_add_watch($inotify, $logDir, IN_CREATE); $fileWatches = []; while (true) { $events = inotify_read($inotify); if ($events) { foreach ($events as $event) { if ($event['mask'] & IN_CREATE && is_file("$logDir/{$event['name']}")) { // Add watch for new log file and track file pointer $fd = fopen("$logDir/{$event['name']}", 'r'); fseek($fd, 0, SEEK_END); $watchId = inotify_add_watch($inotify, "$logDir/{$event['name']}", IN_MODIFY | IN_CLOSE_WRITE); $fileWatches[$watchId] = ['fd' => $fd, 'path' => "$logDir/{$event['name']}"]; } elseif ($event['mask'] & IN_MODIFY && isset($fileWatches[$event['wd']])) { // Read new lines from the file pointer $fd = $fileWatches[$event['wd']]['fd']; while ($line = fgets($fd)) { if (str_contains($line, 'RESUMING MISSION')) { // Run SSH in background using proc_open to avoid blocking $proc = proc_open("ssh -n user@server 'php matlabCmd.php --drone=...'", [], $pipes); proc_close($proc); } } } elseif ($event['mask'] & IN_CLOSE_WRITE && isset($fileWatches[$event['wd']])) { // Clean up watch for closed file fclose($fileWatches[$event['wd']]['fd']); inotify_rm_watch($inotify, $event['wd']); unset($fileWatches[$event['wd']]); } } } usleep(100000); // Reduce CPU usage }
PHP's approach is more scalable and easier to manage for multiple drones.
4. Fixing SSH Background Command Causing Loop Exit
The issue here is that SSH inherits the script's standard input (the inotifywait process pipe) when run in the background. When SSH reads from this pipe, it consumes the inotify output, causing the read loop to receive an EOF and exit.
Solution:
Add the -n flag to SSH to redirect its stdin to /dev/null, and redirect output to avoid clutter:
ssh -n user@another_server "/usr/bin/php /path_to_matlab_command/matlabCmd.php --drone=${vehicle}" > /dev/null 2>&1 &
The -n flag prevents SSH from reading from the script's input pipe, keeping the read loop connected to inotifywait's output.
General Optimizations
- Add logging: Write script activity (events, errors, actions) to a dedicated log file for debugging.
- Handle file rotation: Use
tail -Finstead oftail -fto automatically follow rotated log files. - Clean up stale processes: In bash, add a trap to kill background monitors when the script exits:
trap 'kill $(jobs -p)' EXIT. - Avoid race conditions: When handling file creation, add a short delay or check if the file is fully written before starting monitoring.
内容的提问来源于stack exchange,提问作者Laurent Mortier

