You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用AWK处理文本文件并生成含分类列的新文件

Solution for Adding Categorization Column to TSV File

Got it, let's work through this problem together. You've got a tab-separated (TSV) file with 4 columns, and you need to add a 5th column that tags each row as CA, NON, or BA based on your rules. From what you shared, here's a practical implementation using awk—a tool perfect for text column processing.

Step 1: Clarify the Rules (as provided)

Let's restate the rules to make sure we're aligned:

  • Rule 1: If the value in Column 1 is unique (appears only once in the entire file), mark the 5th column as NON.
  • Rule 2: If the calculation (number after colon in Column 1 + Column 2 value) - number after dash in Column 1 ≤ -31, mark as CA.
  • Default: For all other cases (Column 1 is duplicated, and the calculation doesn't meet Rule 2), mark as BA (since you specified three categories).

Step 2: awk Script Implementation

We'll use a two-pass approach with awk: first to count how many times each Column 1 value appears, then to apply the categorization rules.

Full Script (save as categorize.awk)

BEGIN {
    FS = "\t"  # Set input field separator to tab
    OFS = "\t" # Set output field separator to tab (keeps TSV format)
}

# First pass: Count occurrences of each Column 1 value
NR == FNR {
    count[$1]++
    next
}

# Second pass: Apply categorization rules to each row
{
    # Check Rule 1 first: unique Column 1 = NON
    if (count[$1] == 1) {
        category = "NON"
    } else {
        # Extract numbers from Column 1 (format: [text]:[num1]-[num2])
        split($1, colon_parts, ":")
        split(colon_parts[2], dash_parts, "-")
        num_col1_colon = dash_parts[1] + 0  # Convert to numeric
        num_col1_dash = dash_parts[2] + 0
        num_col2 = $2 + 0

        # Calculate the required value
        calc_result = (num_col1_colon + num_col2) - num_col1_dash

        # Apply Rule 2: if result ≤ -31 = CA, else BA
        category = (calc_result <= -31) ? "CA" : "BA"
    }

    # Print original columns plus the new category column
    print $0, category
}

How to Run

Assume your input file is input.tsv, and you want to save the output to output.tsv:

awk -f categorize.awk input.tsv input.tsv > output.tsv

Or run it as a one-liner without saving the script:

awk 'BEGIN{FS=OFS="\t"} NR==FNR{count[$1]++;next} {if(count[$1]==1){cat="NON"}else{split($1,c,":");split(c[2],d,"-");calc=(d[1]+$2)-d[2];cat=(calc<=-31)?"CA":"BA"} print $0,cat}' input.tsv input.tsv > output.tsv

Notes

  • If Column 1 has a different format than [text]:[num1]-[num2], let me know the exact structure and I can adjust the split logic.
  • If there are additional conditions for the BA category that you didn't mention, we can tweak the script to include those too.

内容的提问来源于stack exchange,提问作者ARM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:01:52