如何实现跨目录文件多列匹配并输出匹配行?现有单列匹配脚本优化
Fixing Multi-Column Matching for Your File Script
Got it, let's adjust your script to match both the first and second columns from for_matching against your *.file files, instead of just a single column. Here's the revised solution, plus a breakdown of how it works:
Revised Script
for i in */*.file; do awk 'FNR==NR {key=$1 FS $2; a[key]=1; next} {key=$1 FS $2; if (key in a) print $0}' for_matching "$i" > "$i.matched" done
What Changed & Why
Let's walk through the improvements from your original script:
- Fixed file order: We now load
for_matchingfirst (instead of the*.file). This lets us store all the key pairs we need to match before scanning the target files. - Multi-column key: Instead of using just
$1as the match key, we combine the first and second columns with$1 FS $2(FS is awk's default field separator, which handles spaces/tabs). This ensures we only match rows where both columns exactly match the entries infor_matching. - Simplified array storage: We only store a flag (
a[key]=1) instead of the full line fromfor_matching—since we just need to check if the key exists, not retrieve the original line. - Output the correct line: We print
$0(the full line from the*.filefile) when a match is found, which aligns with your expected output.
Test with Your Example
If you have:
for_matchingcontaining:carrot 124555A file
produce/file1.filewith:carrot 124555 1 2 6 hair 9922 2 2 2 tree 2223 2 1 2
Running the script will create produce/file1.file.matched with exactly the line you want:
carrot 124555 1 2 6
内容的提问来源于stack exchange,提问作者user3324491
相关产品推荐
相关产品推荐

