如何用简洁awk命令去除第一列重复记录并保留其余列
Great question! Your initial awk approach is already on the right track, and we can definitely trim it down to a more concise one-liner while retaining all the functionality you need (preserving empty lines, keeping first occurrences of the first column, leaving other columns untouched).
Most Concise Version (Handles Empty Lines)
awk '!NF{print;next}!($1 in a){a[$1];print}' file
Even Shorter (If No Empty Lines Need Preserving)
If your input file doesn't have empty lines that you need to keep, you can simplify it to this ultra-clean command:
awk '!($1 in a){a[$1];print}' file
How It Works
- The associative array
aacts as a tracker: every time we encounter a new value in the first column ($1), we add it toaand print the entire line. - For empty lines (where
NF—number of fields—is 0), the first condition!NF{print;next}ensures we print them immediately without checking the array, preserving their place in the output. - Subsequent lines with duplicate first-column values are skipped entirely, since
$1will already exist ina.
This maintains the original line order, keeps all other columns exactly as they are, and efficiently removes duplicates from the first column in a single pass through the file.
内容的提问来源于stack exchange,提问作者OXXO
相关产品推荐
相关产品推荐

