如何用Bash处理CSV文件:统一重复第一列的对应第二列值
Hey there! Since you're new to Bash, let's go with a straightforward, beginner-friendly approach using awk—it's a standard tool for text processing that works perfectly here.
The Core Idea
We'll use an associative array (a simple key-value list) to remember the first occurrence of each value in the first column, along with its corresponding second column value. Then, every time we encounter that same first column value again, we'll replace its second column with the stored value.
Step-by-Step Command
Assume your CSV file is named input.csv. Run this command in your terminal:
awk -F ',' '!seen[$1] { seen[$1] = $2 } { print $1 "," seen[$1] }' input.csv > output.csv
Let's Break This Down
-F ',': Tellsawkto split each line using commas, so$1refers to the first column and$2to the second.!seen[$1] { seen[$1] = $2 }: Checks if we haven't encountered the first column value ($1) before. If it's new, we store its second column value ($2) in theseenarray under the key$1.{ print $1 "," seen[$1] }: For every line, print the first column value followed by the stored second column value (either the original first occurrence or the saved one for duplicates).> output.csv: Saves the modified data to a new file (output.csv) so we don't overwrite your original file.
Test It With Your Example
Let's create a test file with your sample data to verify:
echo -e "a,2\nb,3\na,1\nc,5\nb,2" > test.csv
Run the awk command on test.csv, then check the result:
cat output.csv
You'll get exactly what you wanted:
a,2 b,3 a,2 c,5 b,3
Optional: Use Last Occurrence Instead
If you wanted to use the last occurrence of the first column's corresponding second value instead of the first, just remove the ! from the command:
awk -F ',' '{ seen[$1] = $2 } { print $1 "," seen[$1] }' input.csv > output.csv
This would update the stored value every time we encounter the first column, resulting in the last seen value being used for all duplicates.
内容的提问来源于stack exchange,提问作者Visahan

