求助:使用Awk处理含INV标识数据,完成去重、极值计算与列操作
AWK Solution for Your Processing Requirements
Got it, let's break down your AWK problem and build a solution that hits all your requirements. Here's a complete script, followed by detailed explanations:
BEGIN { # Set input/output field separator (adjust to match your file's delimiter, e.g., " " for spaces) FS = OFS = "\t" } # Only process rows where 5th column is "INV" $5 == "INV" { # Store the first occurrence of the row (excluding 5th and 6th columns) for each unique 1st column value if (!($1 in unique_rows)) { unique_rows[$1] = "" for (i = 1; i <= NF; i++) { if (i != 5 && i != 6) { unique_rows[$1] = unique_rows[$1] (unique_rows[$1] ? OFS : "") $i } } } # Track the minimum and maximum values from the 6th column for each unique 1st column if (!($1 in min_6) || $6 < min_6[$1]) { min_6[$1] = $6 } if (!($1 in max_6) || $6 > max_6[$1]) { max_6[$1] = $6 } } END { # Process each unique 1st column entry to generate the final output for (id in unique_rows) { # Calculate the required difference diff = max_6[id] - min_6[id] + 1 # Split the stored row into an array for easy manipulation split(unique_rows[id], parts, OFS) # Build the final row according to your column requirements: # - Keep columns 1-4 from original # - Insert min(6th) as new column 7, max(6th) as new column 8 # - Insert diff as column 9 and column 11 # - Preserve other original columns (from original 7th onwards) final_row = parts[1] OFS parts[2] OFS parts[3] OFS parts[4] if (parts[5] != "") final_row = final_row OFS parts[5] # Original 7th column (now position 5) if (parts[6] != "") final_row = final_row OFS parts[6] # Original 8th column (now position 6) final_row = final_row OFS min_6[id] OFS max_6[id] OFS diff if (parts[7] != "") final_row = final_row OFS parts[7] # Original 9th column (now position 10) final_row = final_row OFS diff # Add any remaining original columns (from original 10th onwards) for (i = 8; i <= length(parts); i++) { final_row = final_row OFS parts[i] } # Print the processed row print final_row } }
Key Explanations:
Field Separator Setup: The
BEGINblock sets the input/output field separator to tabs (\t). If your file uses spaces or another delimiter, replace this with the correct value (e.g.,FS = OFS = " "for spaces).Filter and Store Unique Rows:
- We only process rows where the 5th column is
"INV". - For each unique 1st column value, we store the first occurrence of the row excluding the 5th and 6th columns (since we need to delete them in the final output).
- We only process rows where the 5th column is
Track Min/Max Values:
- We maintain two associative arrays (
min_6andmax_6) to store the smallest and largest values from the 6th column for each unique 1st column entry.
- We maintain two associative arrays (
Generate Final Output:
- In the
ENDblock, we iterate over each unique 1st column value. - Calculate
diffasmax - min + 1. - We reconstruct the final row by:
- Keeping the original 1st-4th columns.
- Inserting the min/max values in the 7th and 8th column positions (as per your requirement).
- Adding
diffto both the 9th and 11th column positions. - Preserving any other original columns that existed beyond the 6th.
- In the
Adjustments for Your Specific File:
- If your input file doesn't have columns beyond the 6th, you can simplify the final row construction to:
final_row = parts[1] OFS parts[2] OFS parts[3] OFS parts[4] OFS min_6[id] OFS max_6[id] OFS diff OFS "" OFS diff - If the original columns beyond the 6th are in different positions, tweak the
final_rowbuilding logic to match your desired output structure.
内容的提问来源于stack exchange,提问作者as7951
相关产品推荐
相关产品推荐

