You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Linux命令将提取的非行首大写单词生成两两组合?

Generate Unique Pairwise Name Combinations with Linux Commands

Got it, let's tackle this problem step by step. You've already extracted your target names into nodes.csv, and now you need to generate all unique, unordered pairwise combinations (like John, Beatrice or Beatrice, Lucio)—no duplicate edges, no self-matching pairs. Here are two efficient command-line solutions:

Awk is perfect for this task because it’s fast and handles large datasets smoothly. First, we’ll clean up your input file to remove duplicates and empty lines, then generate the combinations.

Step 1: Clean the Input File

Run this to filter out empty lines, remove duplicate entries, and save the result to a cleaned file:

grep -v '^$' nodes.csv | sort -u > unique_nodes.txt
  • grep -v '^$': Removes any blank lines from nodes.csv
  • sort -u: Sorts the names and removes duplicates

Step 2: Generate Pairwise Combinations

Use this Awk command to create all valid unordered pairs:

awk 'NR == FNR { arr[++n] = $0; next } { for (i = 1; i < FNR; i++) print arr[i] ", " $0 }' unique_nodes.txt unique_nodes.txt > edges.csv

How this works:

  • We tell Awk to read unique_nodes.txt twice:
    1. On the first pass (NR == FNR), we store every name in an array called arr.
    2. On the second pass, for each name, we pair it with every name that came before it in the array (i < FNR). This ensures we only get each pair once (e.g., John, Beatrice but not Beatrice, John) and never pair a name with itself.
  • The output is saved directly to edges.csv.

Method 2: Use a Bash Script (For Smaller Datasets)

If your nodes.csv is small, a simple bash loop is easy to read and modify:

Step 1: Clean the Input File (Same as Above)

grep -v '^$' nodes.csv | sort -u > unique_nodes.txt

Step 2: Run the Bash Loop

Create and execute this script (or run it directly in your terminal):

#!/bin/bash
# Read all cleaned names into an array
mapfile -t names < unique_nodes.txt

# Loop through each name, pair it with subsequent names
for ((i=0; i<${#names[@]}; i++)); do
    for ((j=i+1; j<${#names[@]}; j++)); do
        echo "${names[i]}, ${names[j]}"
    done
done > edges.csv
  • mapfile -t names: Reads each line of unique_nodes.txt into the names array
  • The nested loops iterate through the array, pairing each name only with names that come after it (avoiding duplicates and self-pairs)

Verify the Output

You can check the first few lines of your result with:

head edges.csv

This should show combinations exactly like the format you want: John, Beatrice, John, Lucio, etc.

内容的提问来源于stack exchange,提问作者Emanuele Bosimini

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 09:07:49