使用AWK处理文本文件并生成含分类列的新文件
Solution for Adding Categorization Column to TSV File
Got it, let's work through this problem together. You've got a tab-separated (TSV) file with 4 columns, and you need to add a 5th column that tags each row as CA, NON, or BA based on your rules. From what you shared, here's a practical implementation using awk—a tool perfect for text column processing.
Step 1: Clarify the Rules (as provided)
Let's restate the rules to make sure we're aligned:
- Rule 1: If the value in Column 1 is unique (appears only once in the entire file), mark the 5th column as
NON. - Rule 2: If the calculation
(number after colon in Column 1 + Column 2 value) - number after dash in Column 1 ≤ -31, mark asCA. - Default: For all other cases (Column 1 is duplicated, and the calculation doesn't meet Rule 2), mark as
BA(since you specified three categories).
Step 2: awk Script Implementation
We'll use a two-pass approach with awk: first to count how many times each Column 1 value appears, then to apply the categorization rules.
Full Script (save as categorize.awk)
BEGIN { FS = "\t" # Set input field separator to tab OFS = "\t" # Set output field separator to tab (keeps TSV format) } # First pass: Count occurrences of each Column 1 value NR == FNR { count[$1]++ next } # Second pass: Apply categorization rules to each row { # Check Rule 1 first: unique Column 1 = NON if (count[$1] == 1) { category = "NON" } else { # Extract numbers from Column 1 (format: [text]:[num1]-[num2]) split($1, colon_parts, ":") split(colon_parts[2], dash_parts, "-") num_col1_colon = dash_parts[1] + 0 # Convert to numeric num_col1_dash = dash_parts[2] + 0 num_col2 = $2 + 0 # Calculate the required value calc_result = (num_col1_colon + num_col2) - num_col1_dash # Apply Rule 2: if result ≤ -31 = CA, else BA category = (calc_result <= -31) ? "CA" : "BA" } # Print original columns plus the new category column print $0, category }
How to Run
Assume your input file is input.tsv, and you want to save the output to output.tsv:
awk -f categorize.awk input.tsv input.tsv > output.tsv
Or run it as a one-liner without saving the script:
awk 'BEGIN{FS=OFS="\t"} NR==FNR{count[$1]++;next} {if(count[$1]==1){cat="NON"}else{split($1,c,":");split(c[2],d,"-");calc=(d[1]+$2)-d[2];cat=(calc<=-31)?"CA":"BA"} print $0,cat}' input.tsv input.tsv > output.tsv
Notes
- If Column 1 has a different format than
[text]:[num1]-[num2], let me know the exact structure and I can adjust thesplitlogic. - If there are additional conditions for the
BAcategory that you didn't mention, we can tweak the script to include those too.
内容的提问来源于stack exchange,提问作者ARM
相关产品推荐
相关产品推荐

