如何用单个高效Bash脚本实现多字符串匹配日志行分文件输出
Efficient Bash Script to Capture Last Matching Lines per Search String
Got it, let's solve this properly. The biggest win for efficiency here is not scanning the source log multiple times (once per search string) — instead, we'll parse it a single time, track the last matching line for each target string, then write those results to their respective files.
Here's a ready-to-use script that's both efficient and easy to maintain:
#!/bin/bash # -------------------------- Configuration -------------------------- # Define your source log file path SOURCE_LOG="/path/to/your/source.log" # Define your search strings and their corresponding output files # Format: "search_string" -> "output_file_path" declare -A SEARCH_MAP=( ["test_string_1"]="/path/to/output1.log" ["test_string_2"]="/path/to/output2.log" ["ERROR: Connection failed"]="/path/to/connection_errors.log" ["WARN: Disk space low"]="/path/to/disk_warnings.log" ) # ------------------------------------------------------------------- # Initialize an associative array to store the last matching line for each string declare -A LAST_MATCHES # Read the source log line by line (efficient, no subshell overhead) while IFS= read -r line; do # Check each search string against the current line for search_str in "${!SEARCH_MAP[@]}"; do # Use wildcard matching (adjust to =~ for regex if needed) if [[ "$line" == *"$search_str"* ]]; then # Update the last match for this string LAST_MATCHES["$search_str"]="$line" # Optional: Uncomment below to skip checking other strings for this line # break fi done done < "$SOURCE_LOG" # Write the last matches to their respective output files for search_str in "${!SEARCH_MAP[@]}"; do output_file="${SEARCH_MAP[$search_str]}" # Only write if we found a match (remove this check for empty files on no matches) if [[ -n "${LAST_MATCHES[$search_str]}" ]]; then echo "${LAST_MATCHES[$search_str]}" > "$output_file" echo "Wrote last match for '$search_str' to $output_file" else echo "No matches found for '$search_str' — skipping $output_file" fi done
Key Details & Optimizations:
- Single Pass Parsing: The script reads the source log exactly once, which is way faster than looping through the log file for each search string (critical for large log files).
- Associative Arrays: We use Bash associative arrays to cleanly map search strings to both their output files and their last matching lines — this keeps the script organized and easy to extend.
- Flexible Matching: The current script uses simple wildcard matching (
== *"$search_str"*). If you need regex matching, replace that line withif [[ "$line" =~ $search_str ]]; then(regex strings don't need surrounding quotes). - Space-Safe Handling: Both search strings and file paths with spaces work correctly thanks to proper quoting throughout the script.
- Optional Speed Boost: Uncomment the
breakinside the for loop if lines will only ever match one search string — this saves a tiny bit of processing time per line.
How to Use:
- Update the
SOURCE_LOGpath to point to your input log file. - Modify the
SEARCH_MAPassociative array to add/remove your search strings and their corresponding output files. - Make the script executable:
chmod +x last_matches.sh - Run it:
./last_matches.sh
Edge Cases to Consider:
- If a search string has no matches, the script skips writing to its output file (remove the
if [[ -n ... ]]check if you want empty files instead). - For extremely large log files, this script remains efficient because it processes lines one at a time (no loading the entire log into memory).
内容的提问来源于stack exchange,提问作者Colin Davis
相关产品推荐
相关产品推荐

