如何对比两个文件并提取相同行至新文件?awk/grep使用遇阻
I have two files, file1 and file2, and I need to compare them then output the lines from file2 that contain the entire lines from file1 into a new file. I tried using awk and grep commands but didn't get the expected result, even after looking through a lot of related solutions.
File Contents
File 1
UDP-N-acetylenolpyruvoylglucosamine reductase K00075 murB; UDP-N-acetylmuramate dehydrogenase EC:1.3.1.98
UDP-N-acetylglucosamine 1-carboxyvinyltransferase K00790 murA; UDP-N-acetylglucosamine 1-carboxyvinyltransferase EC:2.5.1.7
UDP-N-acetylmuramate--L-alanine ligase K01924 murC; UDP-N-acetylmuramate--alanine ligase EC:6.3.2.8
File 2
D ZPR_2530 UDP-N-acetylglucosamine 1-carboxyvinyltransferase K00790 murA; UDP-N-acetylglucosamine 1-carboxyvinyltransferase EC:2.5.1.7
D ZPR_3743 UDP-N-acetylenolpyruvoylglucosamine reductase K00075 murB; UDP-N-acetylmuramate dehydrogenase EC:1.3.1.98
D ZPR_3807 UDP-N-acetylmuramate--L-alanine ligase K01924 murC; UDP-N-acetylmuramate--alanine ligase EC:6.3.2.8
D ZPR_3810 UDP-N-acetylmuramoylalanine--D-glutamate ligase K01925 murD; UDP-N-acetylmuramoylalanine--D-glutamate ligase EC:6.3.2.9
D ZPR_3812 UDP-N-acetylmuramoylalanyl-D-glutamate--2 K01928 murE; UDP-N-acetylmuramoyl-L-alanyl-D-glutamate--2,6-diaminopimelate ligase EC:6.3.2.13
D ZPR_0820 D-alanyl-alanine synthetase A K01921 ddl; D-alanine-D-alanine ligase EC:6.3.2.4
D ZPR_3928 UDP-N-acetylmuramoyl-tripeptide--D-alanyl-D-alanine ligase K01929 murF; UDP-N-acetylmuramoyl-tripeptide--D-alanyl-D-alanine ligase EC:6.3.2.10
D ZPR_4441 putative undecaprenol kinase K06153 bacA; undecaprenyl-diphosphatase EC:3.6.1.27
D ZPR_3043 PAP2 superfamily membrane protein K19302 bcrC; undecaprenyl-diphosphatase EC:3.6.1.27
Expected Output
D ZPR_3743 UDP-N-acetylenolpyruvoylglucosamine reductase K00075 murB; UDP-N-acetylmuramate dehydrogenase EC:1.3.1.98
D ZPR_2530 UDP-N-acetylglucosamine 1-carboxyvinyltransferase K00790 murA; UDP-N-acetylglucosamine 1-carboxyvinyltransferase EC:2.5.1.7
D ZPR_3807 UDP-N-acetylmuramate--L-alanine ligase K01924 murC; UDP-N-acetylmuramate--alanine ligase EC:6.3.2.8
Commands I've Tried
awk 'NR==FNR{a[$1]=$NF;next;} {print ($0 ? a[$1] OFS $0 :$0)}' no-tab-file.txt Non-homo-Dzpr00001.txtgrep -Ff Non-homo-Dzpr00001.txt no-tab-file.txt
Solutions That Work
Using grep
The main issue with your existing grep command was the order of the files. The -f flag tells grep to use lines from the first file as patterns to search for in the second file—you had them reversed, so it was searching for lines from file2 in file1 instead of the other way around.
Here's the corrected command:
grep -Ff file1 file2 > output.txt
-Fensures grep treats each line fromfile1as a fixed string (not a regex), which is crucial because your lines contain special characters like;and:that would be misinterpreted as regex syntax otherwise.- This command will output all lines from
file2that contain any full line fromfile1, which exactly matches your expected result.
Using awk
Your original awk script was storing just the first field of file1 and mapping it to the last field, which isn't aligned with what you need. Instead, we can store all lines from file1 in an array, then check each line in file2 to see if it contains any of those stored lines.
Basic implementation:
awk 'NR==FNR { stored_lines[$0]; next } { for (line in stored_lines) if ($0 ~ line) print $0 }' file1 file2 > output.txt
NR==FNR { stored_lines[$0]; next }: Reads every line fromfile1into thestored_linesarray (using the entire line as the key).for (line in stored_lines) if ($0 ~ line) print $0: For each line infile2, checks if it contains any line fromfile1; if yes, prints thefile2line.
More precise implementation (check if file1 line is a suffix):
If you want to make absolutely sure the match is at the end of the file2 line (since that's how your data is structured), you can use this version:
awk 'NR==FNR { stored_lines[$0]; len[$0] = length($0); next } { for (line in stored_lines) if (substr($0, length($0) - len[line] + 1) == line) print $0 }' file1 file2 > output.txt
This checks if the end of the file2 line exactly matches the entire line from file1, avoiding any accidental partial matches elsewhere in the line.
内容的提问来源于stack exchange,提问作者Alina

