Linux下多文件匹配行识别的更优方案问询
comm Great question—using chained comm calls like comm -12 <(sort file1.txt) <(sort file2.txt) | comm -12 - <(sort file3.txt) works, but it’s definitely not readable once you have three or more files. Here are a few cleaner, ready-to-use solutions so you don’t have to reinvent the wheel entirely:
1. One-Liner with sort + uniq + awk (Simplest Approach)
This method works for any number of files and is far more concise than chained comm. The idea is to count how many times each line appears across all files—lines that show up exactly once per file are the common ones:
sort file*.txt | uniq -c | awk -v n=$(ls file*.txt | wc -l) '$1 == n {print $2}'
Breakdown:
sort file*.txt: Merges and sorts all input files into a single stream.uniq -c: Counts occurrences of each unique line.awk ...: Filters lines where the count matches the total number of input files (meaning the line exists in every file).
If you want to specify files explicitly instead of using a wildcard, replace file*.txt with your list, and adjust the n value accordingly:
sort file1.txt file2.txt file3.txt | uniq -c | awk -v n=3 '$1 == n {print $2}'
2. awk Script for Unordered Files (No Pre-Sorting)
If your files are large or you don’t want to sort them first, use this awk script to track which lines exist in every file. It avoids sorting but uses more memory since it stores all unique lines:
awk ' BEGIN { # Read all lines from each input file and track which files they appear in for (i = 1; i < ARGC; i++) { while ((getline line < ARGV[i]) > 0) { seen[line][ARGV[i]] = 1 } close(ARGV[i]) } ARGC = 1 # Prevent awk from processing files again } END { # Check which lines exist in all files for (line in seen) { count = 0 for (file in seen[line]) count++ if (count == ARGC-1) print line } } ' file1.txt file2.txt file3.txt
3. Bash Function to Wrap Chained comm (Familiar Logic)
If you prefer sticking with comm but want better readability, wrap the logic in a bash function. This hides the messy nesting and lets you call it with a simple list of files:
multi_comm() { if [ $# -lt 2 ]; then echo "Usage: multi_comm file1 file2 [file3 ...]" >&2 return 1 fi # Start with the sorted first file current=$(sort "$1") shift # Iterate through remaining files, narrowing down common lines for file in "$@"; do current=$(comm -12 <(echo "$current") <(sort "$file")) done echo "$current" }
Use it like this:
multi_comm file1.txt file2.txt file3.txt
4. Interactive Script (If You Want User Input)
If you want a fully interactive tool that prompts for filenames (like you mentioned), here’s a simple script that combines the sort+uniq+awk method with input validation:
#!/bin/bash echo "Enter the filenames to find common lines in (separate with spaces):" read -a input_files # Validate all files exist for file in "${input_files[@]}"; do if [ ! -f "$file" ]; then echo "Error: File '$file' does not exist!" >&2 exit 1 fi done if [ "${#input_files[@]}" -lt 2 ]; then echo "Error: You need at least two files!" >&2 exit 1 fi # Find common lines sort "${input_files[@]}" | uniq -c | awk -v n="${#input_files[@]}" '$1 == n {print $2}'
Save this as find_common_lines.sh, make it executable (chmod +x find_common_lines.sh), and run it—you’ll get a prompt to enter your filenames.
All these options avoid the messy chained comm syntax while getting the job done. Pick the one that fits your workflow best!
内容的提问来源于stack exchange,提问作者Yusif_Nurizade

