使用awk基于公共列(ID列)合并两个文本文件
To merge your two files based on the 3rd column (ID) while retaining File2's format and only keeping rows present in both files, you can use a clean awk command that efficiently checks for matching IDs.
Here's the command to run:
awk 'NR==FNR {ids[$3]++; next} /^#/ {print; next} ids[$3]' file1.txt file2.txt > merged_output.txt
How this works:
- Track IDs from File1: The
NR==FNRcondition only triggers when processing the first input file (file1.txt). We store every ID (3rd column,$3) in an associative arrayidsto keep a record of which IDs we need to retain. - Preserve File2's header: For the second file (
file2.txt), if the line starts with#(the header), we print it immediately to maintain your desired output structure. - Filter matching rows: For non-header lines in File2, we check if the ID exists in our
idsarray. If it does, we print the entire line (including all columns from File2, exactly as you specified).
Example Output:
Running this command on your sample files will produce:
#CHROM POS ID REF ALT FORMAT 1 61442 rs2531261 A G GT:GL:GQ:DP:CGA_RDP:EHQ:HQ:CGA_CEGL:PS:CGA_CEHQ:AD 1 924629 rs28622096 A G,T GT:GL:GQ:DP:CGA_RDP:EHQ:HQ:CGA_CEGL:PS:CGA_CEHQ:AD 1 974662 rs2465135 G T GT:GL:GQ:DP:CGA_RDP:EHQ:HQ:CGA_CEGL:PS:CGA_CEHQ:AD 1 995481 rs9442393 T G GT:GL:GQ:DP:CGA_RDP:EHQ:HQ:CGA_CEGL:PS:CGA_CEHQ:AD
This excludes the row from File2 that doesn't appear in File1 (rs34235844) and matches your desired output perfectly.
Alternative: Using join (requires sorted files)
If you prefer using join, you'll first need to sort both files by the 3rd column (while preserving headers), then join them. Here's how:
# Sort File1 (preserve header) head -1 file1.txt > sorted_file1.txt tail -n +2 file1.txt | sort -k3,3 >> sorted_file1.txt # Sort File2 (preserve header) head -1 file2.txt > sorted_file2.txt tail -n +2 file2.txt | sort -k3,3 >> sorted_file2.txt # Join on 3rd column, output all columns from File2 join -1 3 -2 3 -o 2.1,2.2,2.3,2.4,2.5,2.6 sorted_file1.txt sorted_file2.txt > merged_output.txt
This approach works reliably but requires sorting, which may be slower for very large files compared to the awk method.
内容的提问来源于stack exchange,提问作者June
相关产品推荐
相关产品推荐

