You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

批量处理dlg文件提取CLUSTERING HISTOGRAM最大簇行的Bash脚本需求

Bash Script to Extract Largest Cluster from DLG Files

Here's a complete, polished script that replaces your existing echo/grep logic and delivers exactly the functionality you need:

#!/bin/bash

# Initialize output file (clear existing content if present)
OUTPUT_FILE="$PWD/final_log.txt"
> "$OUTPUT_FILE"

# Iterate over all .dlg files in the current working directory
for f in *.dlg; do
    # Skip loop if no .dlg files exist (avoids error messages)
    [ -e "$f" ] || continue

    echo "Processing $f..."
    file_name2=$(basename "$f")
    file_name="${file_name2%.dlg}"  # Correct syntax to remove .dlg suffix

    # Use awk to find the row with the most # in the Histogram column
    largest_cluster_row=$(awk '
        BEGIN { max_hashes = 0; result_row = "" }
        # Trigger table tracking when we hit the header
        /CLUSTERING HISTOGRAM/ { in_table = 1; next }
        # Process only data rows (start with a digit) while in the table
        in_table && /^[[:digit:]]/ {
            # Count number of # in the last field (Histogram column)
            current_hashes = length($NF) - length(gsub("#", "", $NF))
            # Update result if current row has more # than previous max
            if (current_hashes > max_hashes) {
                max_hashes = current_hashes
                result_row = $0
            }
        }
        # Stop tracking table when we hit the bottom separator line
        in_table && /^_____/ { in_table = 0 }
        # Print the final result row at the end
        END { print result_row }
    ' "$f")

    # Write to output file in your specified format
    echo "\"$file_name\": $largest_cluster_row" >> "$OUTPUT_FILE"
done

echo "Processing complete! Results saved to $OUTPUT_FILE"

Key Details Explained:

  • Output File Setup: We clear final_log.txt at the start to ensure fresh results every time you run the script.
  • Safe File Iteration: The [ -e "$f" ] || continue line prevents errors if there are no .dlg files in the directory.
  • Awk Logic Breakdown:
    • We set a flag (in_table) when we encounter the CLUSTERING HISTOGRAM header, so we know when we're inside the target table.
    • We only process rows that start with a digit (actual data rows), skipping headers and separator lines.
    • For each data row, we count the number of # in the last column by comparing the field length before and after removing all # characters.
    • We keep track of the row with the highest # count. If multiple rows have the same max count, this script picks the first one encountered (change > to >= in the condition if you want the last matching row instead).
    • We turn off the in_table flag when we hit the bottom separator line (starting with _____), since the table ends there.
  • Exact Output Format: The script writes entries in the exact format you provided, with quoted filenames and the full matching row.

Example Output:

Your final_log.txt will look identical to your sample:

"Name of the file 1": 3 | -5.47 | 17 | -5.44 | 2 |##
"Name_of_the_file_2": 1 | -5.99 | 13 | -5.98 | 16 |################
"Name_of_the_file_3": 2 | -4.78 | 19 | -4.44 | 3 |###

内容的提问来源于stack exchange,提问作者user3470313

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 08:32:28