为何sort -u命令会遗漏文件中的部分唯一行?
背景
假设有一个名为main.txt的文件,内容如下:
line 1 䔍 䏝 line 4 line 5 䏝
该文件包含5条唯一行,其中第3行和第5行包含相同字符:"䏝"(U+43DD)。
执行sort -u main.txt命令后,得到如下结果:
䔍 line 1 line 4 line 5 䏝
可以看到,内容为䏝的第3行虽为唯一行,却未出现在结果中。
问题
为何sort -u命令未将第3行纳入结果,尽管它是唯一行?
系统信息
file命令输出
执行file main.txt的结果:
file main.txt main.txt: Unicode text, UTF-8 text
locale设置
系统locale命令的输出:
locale LANG=en_US.UTF-8 LC_CTYPE="en_US.UTF-8" LC_NUMERIC="en_US.UTF-8" LC_TIME="en_US.UTF-8" LC_COLLATE="en_US.UTF-8" LC_MONETARY="en_US.UTF-8" LC_MESSAGES="en_US.UTF-8" LC_PAPER="en_US.UTF-8" LC_NAME="en_US.UTF-8" LC_ADDRESS="en_US.UTF-8" LC_TELEPHONE="en_US.UTF-8" LC_MEASUREMENT="en_US.UTF-8" LC_IDENTIFICATION="en_US.UTF-8" LC_ALL=
sort版本
sort --version sort (GNU coreutils) 9.1 Copyright (C) 2022 Free Software Foundation, Inc. License GPLv3+: GNU GPL version 3 or later <https://gnu.org/licenses/gpl.html>. This is free software: you are free to change and redistribute it. There is NO WARRANTY, to the extent permitted by law. Written by Mike Haertel and Paul Eggert.
包管理器信息
查询sort所属包:
pacman -Q -o "$(which sort)" /usr/bin/sort is owned by coreutils 9.1-1
本地已安装coreutils包详情:
pacman -Q -i coreutils Name : coreutils Version : 9.1-1 Description : The basic file, shell and text manipulation utilities of the GNU operating system Architecture : x86_64 URL : https://www.gnu.org/software/coreutils/ Licenses : GPL3 Groups : None Provides : None Depends On : glibc acl attr gmp libcap openssl Optional Deps : None Required By : base ca-certificates-utils dkms java-runtime-common linux mkinitcpio p11-kit pacman util-linux Optional For : usbutils Conflicts With : None Replaces : None Installed Size : 15.24 MiB Packager : Sébastien Luttringer <seblu@seblu.net> Build Date : Sun 17 Apr 2022 01:21:13 PM -05 Install Date : Thu 28 Apr 2022 12:14:56 AM -05 Install Reason : Installed as a dependency for another package Install Script : No Validated By : Signature
问题解答
GNU sort -u的去重逻辑不是基于字符串完全相等,而是看当前locale排序规则下两行是否被判定为排序等价。只要排序时两行被视为“同一排序位置的元素”,-u就会将它们当作重复项,只保留其中一行。
具体到这个案例:
字符䏝(U+43DD)属于CJK统一表意文字扩展A区,在en_US.UTF-8的排序规则中,单独的䏝和字符串line 5 䏝被判定为排序等价——sort认为这两行的排序键相同,因此-u直接过滤掉了单独的䏝那一行。
如果需要按严格的字符串相等去重,可以强制使用C locale(基于字节的原始比较逻辑):
LC_COLLATE=C sort -u main.txt
执行该命令后,所有5行都会被保留,因为C locale完全按字节序列比较,这两行的字节内容完全不同,不会被判定为重复。
内容的提问来源于stack exchange,提问作者Rodrigo Morales
相关产品推荐
相关产品推荐

