为何Bash sort命令的-n与-u选项组合会过滤掉IP地址69.156.7.43?
sort -n and -u remove the IP 69.156.7.43? Let's break down exactly what's happening here—the issue boils down to how sort interprets numerical sorting (-n) for IP addresses, and how the -u (unique) option interacts with that behavior.
1. The core problem: sort -n doesn't handle IP addresses correctly by default
When you use sort -n with IP addresses formatted as a.b.c.d, sort treats the entire string as a floating-point number, not a sequence of four separate numerical segments.
For your two IPs:
69.156.151.245gets parsed as the float69.156(sincesortstops interpreting at the second dot—floats only allow one decimal point)69.156.7.43also gets parsed as69.156(same reason: the second dot truncates the rest of the string)
To sort -n, these two IPs are numerically identical. That's the root cause!
2. How -u interacts with -n
The -u option removes duplicate lines based on how sort defines equality under the current sorting rules. When combined with -n, sort sees 69.156.151.245 and 69.156.7.43 as identical, so it keeps only one of them (whichever comes last in the sorted input stream—in your case, 69.156.151.245) and discards the other.
Let's tie this to your test cases:
- Without
-n(lexicographical sort):sortcompares each character individually.69.156.151.245and69.156.7.43are lex distinct (the third segment starts with1vs7), so-udoesn't remove either. - Without
-u(numerical sort):sortgroups the two "equal" IPs together, but since-uisn't present, all instances (including duplicates of69.156.7.43) are retained.
3. Why -V fixes it
The -V (version sort) option is built specifically for strings like IP addresses, version numbers, etc., where segments are separated by dots. It treats each dot-separated part as an independent numerical value, so:
- It compares the first segment (69 vs 69: equal)
- Then the second (156 vs 156: equal)
- Then the third (7 vs 151: 7 is smaller, so
69.156.7.43comes first) - Finally the fourth segment if needed
When combined with -u, -V correctly identifies unique IPs because it compares each segment properly—no more treating entire IPs as a single float.
Alternative fix if you prefer -n
Your working command with -V is already the cleanest approach, but if you wanted to use -n correctly, you'd need to explicitly apply numerical sorting to each of the four segments:
| cut -w -f11 | sort -t. -k1,1n -k2,2n -k3,3n -k4,4n -u
This tells sort to sort numerically on each dot-separated field individually, which will also correctly handle IPs and deduplicate them properly.
内容的提问来源于stack exchange,提问作者James B. Byrne

