如何用awk处理TSV文件:合并同前缀行并求和第二列?
Here's a straightforward awk script that handles your requirement perfectly. It uses an associative array to accumulate sums for each base ID, regardless of whether the line is the original or the _2 suffix version.
BEGIN { FS = "\t"; OFS = "\t" } { # Get the base ID by stripping "_2" if present key = $1 if (key ~ /_2$/) { key = substr(key, 1, length(key)-2) } # Accumulate the sum for this base ID sum[key] += $2 } END { # Print all base IDs and their total sums for (k in sum) { print k, sum[k] } }
How to Use
Run it against your TSV file like this:
awk -f merge_tsv.awk your_file.tsv
Or directly in the command line:
awk 'BEGIN { FS = "\t"; OFS = "\t" } { key = $1; if (key ~ /_2$/) key = substr(key,1,length(key)-2); sum[key] += $2 } END { for(k in sum) print k, sum[k] }' your_file.tsv
Example
Input TSV:
geneA 25 geneA_2 15 geneB 30 geneC 10 geneC_2 5
Output:
geneA 40 geneB 30 geneC 15
Breakdown of the Script
BEGINBlock: Sets input (FS) and output (OFS) field separators to tab, which is critical for correctly parsing TSV files.- Main Processing Block:
- Takes the first field as the initial key.
- Checks if the key ends with
_2using a regex match (/_2$/). If yes, trims those two characters to get the base ID. - Adds the second field's value to the sum stored in the
sumarray under the base ID key.
ENDBlock: Loops through the associative array and prints each base ID along with its total sum in TSV format.
Key Notes
- Order Doesn't Matter: Whether the base ID line comes before or after the
_2line, the sum will be accurate. - Handles Lone Entries: Lines without a corresponding
_2version will just keep their original value in the output. - Efficient for Large Files: Awk processes lines one at a time, so it won't bog down even with huge TSV files.
内容的提问来源于stack exchange,提问作者john
相关产品推荐
相关产品推荐

