如何仅去除第一列重复值并将第二列数据横向排列?
需求与解决方案
需求
仅去除第一列的重复值,将同一第一列对应的第二列数据横向排列,每行以第一列的值作为首个元素,同时要求各列对齐、行间距清晰便于查看。
输入数据
blunt-snouted_clingfish rdh14b blunt-snouted_clingfish LOC114457682 blunt-snouted_clingfish rngtt blunt-snouted_clingfish cnr1 blunt-snouted_clingfish akirin2 blunt-snouted_clingfish rars2 blunt-snouted_clingfish slc35a1 blunt-snouted_clingfish LOC114457715 blunt-snouted_clingfish rhag Chinese_tongue_sole nt5c1bb Chinese_tongue_sole si:dkey-174m14.3 Chinese_tongue_sole rdh14b Chinese_tongue_sole LOC103381225 Chinese_tongue_sole rngtt Chinese_tongue_sole cnr1 Chinese_tongue_sole akirin2 Chinese_tongue_sole rars2 Chinese_tongue_sole riox1 Chinese_tongue_sole ndufb1 Chinese_tongue_sole cpsf2 Helicophagus_hypophthalmus_Sauvage,_1878 myo6a Helicophagus_hypophthalmus_Sauvage,_1878 LOC113528782 Helicophagus_hypophthalmus_Sauvage,_1878 mei4 Helicophagus_hypophthalmus_Sauvage,_1878 nt5e Helicophagus_hypophthalmus_Sauvage,_1878 snx14 Helicophagus_hypophthalmus_Sauvage,_1878 cnr1 Helicophagus_hypophthalmus_Sauvage,_1878 rngtt Helicophagus_hypophthalmus_Sauvage,_1878 pnrc1 Helicophagus_hypophthalmus_Sauvage,_1878 LOC113528790 Helicophagus_hypophthalmus_Sauvage,_1878 LOC113529170 Helicophagus_hypophthalmus_Sauvage,_1878 c30h8orf82
期望输出
blunt-snouted_clingfish rdh14b LOC114457682 rngtt cnr1 akirin2 rars2 slc35a1 LOC114457715 rhag Chinese_tongue_sole nt5c1bb si:dkey-174m14.3 rdh14b LOC103381225 rngtt cnr1 akirin2 rars2 riox1 ndufb1 cpsf2 Helicophagus_hypophthalmus_Sauvage,_1878 myo6a LOC113528782 mei4 nt5e snx14 cnr1 rngtt pnrc1 LOC113528790 LOC113529170 c30h8orf82
尝试过的命令
之前使用的awk命令仅保留了第一列唯一值及对应首个第二列值,不符合需求:
awk '{if (!seen[$1]++) {print $1, $2}}'
解决方案
1. 合并同一第一列的第二列数据
使用以下awk命令,先将同一第一列对应的第二列数据累积,最后统一输出:
awk 'NF {arr[$1] = arr[$1] ? arr[$1] " " $2 : $1 " " $2} END {for (key in arr) print arr[key]}' input.txt
NF:跳过空行,只处理有内容的行arr[$1] = arr[$1] ? arr[$1] " " $2 : $1 " " $2:如果当前第一列已在数组中,就将第二列追加到对应字符串后;如果是首次出现,就初始化字符串为第一列加第二列END块:遍历数组输出所有合并完成的行
2. 实现列对齐
如果需要让输出的各列对齐,可将上述命令的结果通过column命令处理:
awk 'NF {arr[$1] = arr[$1] ? arr[$1] " " $2 : $1 " " $2} END {for (key in arr) print arr[key]}' input.txt | column -t
column -t:自动识别列分隔符并对齐所有列,让输出更易读
内容的提问来源于stack exchange,提问作者Rohan Nath
相关产品推荐
相关产品推荐

