批量处理dlg文件提取CLUSTERING HISTOGRAM最大簇行的Bash脚本需求
Bash Script to Extract Largest Cluster from DLG Files
Here's a complete, polished script that replaces your existing echo/grep logic and delivers exactly the functionality you need:
#!/bin/bash # Initialize output file (clear existing content if present) OUTPUT_FILE="$PWD/final_log.txt" > "$OUTPUT_FILE" # Iterate over all .dlg files in the current working directory for f in *.dlg; do # Skip loop if no .dlg files exist (avoids error messages) [ -e "$f" ] || continue echo "Processing $f..." file_name2=$(basename "$f") file_name="${file_name2%.dlg}" # Correct syntax to remove .dlg suffix # Use awk to find the row with the most # in the Histogram column largest_cluster_row=$(awk ' BEGIN { max_hashes = 0; result_row = "" } # Trigger table tracking when we hit the header /CLUSTERING HISTOGRAM/ { in_table = 1; next } # Process only data rows (start with a digit) while in the table in_table && /^[[:digit:]]/ { # Count number of # in the last field (Histogram column) current_hashes = length($NF) - length(gsub("#", "", $NF)) # Update result if current row has more # than previous max if (current_hashes > max_hashes) { max_hashes = current_hashes result_row = $0 } } # Stop tracking table when we hit the bottom separator line in_table && /^_____/ { in_table = 0 } # Print the final result row at the end END { print result_row } ' "$f") # Write to output file in your specified format echo "\"$file_name\": $largest_cluster_row" >> "$OUTPUT_FILE" done echo "Processing complete! Results saved to $OUTPUT_FILE"
Key Details Explained:
- Output File Setup: We clear
final_log.txtat the start to ensure fresh results every time you run the script. - Safe File Iteration: The
[ -e "$f" ] || continueline prevents errors if there are no.dlgfiles in the directory. - Awk Logic Breakdown:
- We set a flag (
in_table) when we encounter theCLUSTERING HISTOGRAMheader, so we know when we're inside the target table. - We only process rows that start with a digit (actual data rows), skipping headers and separator lines.
- For each data row, we count the number of
#in the last column by comparing the field length before and after removing all#characters. - We keep track of the row with the highest
#count. If multiple rows have the same max count, this script picks the first one encountered (change>to>=in the condition if you want the last matching row instead). - We turn off the
in_tableflag when we hit the bottom separator line (starting with_____), since the table ends there.
- We set a flag (
- Exact Output Format: The script writes entries in the exact format you provided, with quoted filenames and the full matching row.
Example Output:
Your final_log.txt will look identical to your sample:
"Name of the file 1": 3 | -5.47 | 17 | -5.44 | 2 |## "Name_of_the_file_2": 1 | -5.99 | 13 | -5.98 | 16 |################ "Name_of_the_file_3": 2 | -4.78 | 19 | -4.44 | 3 |###
内容的提问来源于stack exchange,提问作者user3470313
相关产品推荐
相关产品推荐

