You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对比两个文件并提取相同行至新文件?awk/grep使用遇阻

Extract matching lines between two files where file1 lines are substrings of file2 lines

I have two files, file1 and file2, and I need to compare them then output the lines from file2 that contain the entire lines from file1 into a new file. I tried using awk and grep commands but didn't get the expected result, even after looking through a lot of related solutions.

File Contents

File 1

UDP-N-acetylenolpyruvoylglucosamine reductase K00075 murB; UDP-N-acetylmuramate dehydrogenase EC:1.3.1.98
UDP-N-acetylglucosamine 1-carboxyvinyltransferase K00790 murA; UDP-N-acetylglucosamine 1-carboxyvinyltransferase EC:2.5.1.7
UDP-N-acetylmuramate--L-alanine ligase K01924 murC; UDP-N-acetylmuramate--alanine ligase EC:6.3.2.8

File 2

D ZPR_2530 UDP-N-acetylglucosamine 1-carboxyvinyltransferase K00790 murA; UDP-N-acetylglucosamine 1-carboxyvinyltransferase EC:2.5.1.7
D ZPR_3743 UDP-N-acetylenolpyruvoylglucosamine reductase K00075 murB; UDP-N-acetylmuramate dehydrogenase EC:1.3.1.98
D ZPR_3807 UDP-N-acetylmuramate--L-alanine ligase K01924 murC; UDP-N-acetylmuramate--alanine ligase EC:6.3.2.8
D ZPR_3810 UDP-N-acetylmuramoylalanine--D-glutamate ligase K01925 murD; UDP-N-acetylmuramoylalanine--D-glutamate ligase EC:6.3.2.9
D ZPR_3812 UDP-N-acetylmuramoylalanyl-D-glutamate--2 K01928 murE; UDP-N-acetylmuramoyl-L-alanyl-D-glutamate--2,6-diaminopimelate ligase EC:6.3.2.13
D ZPR_0820 D-alanyl-alanine synthetase A K01921 ddl; D-alanine-D-alanine ligase EC:6.3.2.4
D ZPR_3928 UDP-N-acetylmuramoyl-tripeptide--D-alanyl-D-alanine ligase K01929 murF; UDP-N-acetylmuramoyl-tripeptide--D-alanyl-D-alanine ligase EC:6.3.2.10
D ZPR_4441 putative undecaprenol kinase K06153 bacA; undecaprenyl-diphosphatase EC:3.6.1.27
D ZPR_3043 PAP2 superfamily membrane protein K19302 bcrC; undecaprenyl-diphosphatase EC:3.6.1.27

Expected Output

D ZPR_3743 UDP-N-acetylenolpyruvoylglucosamine reductase K00075 murB; UDP-N-acetylmuramate dehydrogenase EC:1.3.1.98
D ZPR_2530 UDP-N-acetylglucosamine 1-carboxyvinyltransferase K00790 murA; UDP-N-acetylglucosamine 1-carboxyvinyltransferase EC:2.5.1.7
D ZPR_3807 UDP-N-acetylmuramate--L-alanine ligase K01924 murC; UDP-N-acetylmuramate--alanine ligase EC:6.3.2.8

Commands I've Tried

  • awk 'NR==FNR{a[$1]=$NF;next;} {print ($0 ? a[$1] OFS $0 :$0)}' no-tab-file.txt Non-homo-Dzpr00001.txt
  • grep -Ff Non-homo-Dzpr00001.txt no-tab-file.txt

Solutions That Work

Using grep

The main issue with your existing grep command was the order of the files. The -f flag tells grep to use lines from the first file as patterns to search for in the second file—you had them reversed, so it was searching for lines from file2 in file1 instead of the other way around.

Here's the corrected command:

grep -Ff file1 file2 > output.txt
  • -F ensures grep treats each line from file1 as a fixed string (not a regex), which is crucial because your lines contain special characters like ; and : that would be misinterpreted as regex syntax otherwise.
  • This command will output all lines from file2 that contain any full line from file1, which exactly matches your expected result.

Using awk

Your original awk script was storing just the first field of file1 and mapping it to the last field, which isn't aligned with what you need. Instead, we can store all lines from file1 in an array, then check each line in file2 to see if it contains any of those stored lines.

Basic implementation:

awk 'NR==FNR { stored_lines[$0]; next } { for (line in stored_lines) if ($0 ~ line) print $0 }' file1 file2 > output.txt
  • NR==FNR { stored_lines[$0]; next }: Reads every line from file1 into the stored_lines array (using the entire line as the key).
  • for (line in stored_lines) if ($0 ~ line) print $0: For each line in file2, checks if it contains any line from file1; if yes, prints the file2 line.

More precise implementation (check if file1 line is a suffix):

If you want to make absolutely sure the match is at the end of the file2 line (since that's how your data is structured), you can use this version:

awk 'NR==FNR { stored_lines[$0]; len[$0] = length($0); next } { for (line in stored_lines) if (substr($0, length($0) - len[line] + 1) == line) print $0 }' file1 file2 > output.txt

This checks if the end of the file2 line exactly matches the entire line from file1, avoiding any accidental partial matches elsewhere in the line.


内容的提问来源于stack exchange,提问作者Alina

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:09:52