如何通过Linux命令将提取的非行首大写单词生成两两组合?
Got it, let's tackle this problem step by step. You've already extracted your target names into nodes.csv, and now you need to generate all unique, unordered pairwise combinations (like John, Beatrice or Beatrice, Lucio)—no duplicate edges, no self-matching pairs. Here are two efficient command-line solutions:
Method 1: Use Awk (Recommended for Large Files)
Awk is perfect for this task because it’s fast and handles large datasets smoothly. First, we’ll clean up your input file to remove duplicates and empty lines, then generate the combinations.
Step 1: Clean the Input File
Run this to filter out empty lines, remove duplicate entries, and save the result to a cleaned file:
grep -v '^$' nodes.csv | sort -u > unique_nodes.txt
grep -v '^$': Removes any blank lines fromnodes.csvsort -u: Sorts the names and removes duplicates
Step 2: Generate Pairwise Combinations
Use this Awk command to create all valid unordered pairs:
awk 'NR == FNR { arr[++n] = $0; next } { for (i = 1; i < FNR; i++) print arr[i] ", " $0 }' unique_nodes.txt unique_nodes.txt > edges.csv
How this works:
- We tell Awk to read
unique_nodes.txttwice:- On the first pass (
NR == FNR), we store every name in an array calledarr. - On the second pass, for each name, we pair it with every name that came before it in the array (
i < FNR). This ensures we only get each pair once (e.g., John, Beatrice but not Beatrice, John) and never pair a name with itself.
- On the first pass (
- The output is saved directly to
edges.csv.
Method 2: Use a Bash Script (For Smaller Datasets)
If your nodes.csv is small, a simple bash loop is easy to read and modify:
Step 1: Clean the Input File (Same as Above)
grep -v '^$' nodes.csv | sort -u > unique_nodes.txt
Step 2: Run the Bash Loop
Create and execute this script (or run it directly in your terminal):
#!/bin/bash # Read all cleaned names into an array mapfile -t names < unique_nodes.txt # Loop through each name, pair it with subsequent names for ((i=0; i<${#names[@]}; i++)); do for ((j=i+1; j<${#names[@]}; j++)); do echo "${names[i]}, ${names[j]}" done done > edges.csv
mapfile -t names: Reads each line ofunique_nodes.txtinto thenamesarray- The nested loops iterate through the array, pairing each name only with names that come after it (avoiding duplicates and self-pairs)
Verify the Output
You can check the first few lines of your result with:
head edges.csv
This should show combinations exactly like the format you want: John, Beatrice, John, Lucio, etc.
内容的提问来源于stack exchange,提问作者Emanuele Bosimini

