为何Perl版split /\s+/实现wc工具的单词统计结果异常?
Perl版wc工具单词数统计异常的原因与修复
问题现象
用Perl和Python分别实现Unix的wc字数统计工具,处理约2万行的C源码文件时,Perl版本统计的单词数远高于标准wc工具和Python版本:
实现代码
Perl版本
#!/usr/bin/perl -w use strict; my $lines=0; my $words=0; my $bytes=0; while (<>) { $lines++; $bytes += length; #chomp; # chomp does not make a difference $words += split /\s+/; # split on white spaces } print "$lines $words $bytes\n";
Python版本
#!/usr/bin/python3 import fileinput lines = 0 words = 0 bytes = 0 for line in fileinput.input(): lines += 1 bytes += len(line) words += len(line.strip().split()) # Python split on white spaces by default print(f"{lines} {words} {bytes}")
输出结果
OUTPUT: cat source.c | ./wc.pl 19681 62506 660235 cat source.c | ./pywc.py 19681 46643 660235 cat source.c | wc 19681 46643 660235
原因分析
核心差异在于split处理空白字符的行为:
- 标准wc与Python:单词定义为「由空白字符分隔的非空白字符序列」,会自动忽略开头、结尾的空白,连续空白视为单个分隔符,不会产生空元素。Python的
split()(无参数)默认就是这种行为,搭配strip()完全匹配标准wc的逻辑。 - Perl的
split /\s+/:当字符串开头是空白时,会将开头空白前的空内容作为第一个元素返回。例如一行是int main(),split /\s+/会得到('', 'int', 'main()'),长度为3,比实际单词数多1。大量行累加后,就导致总单词数明显偏高。
修复方案
方案1:使用Perl默认的split行为
Perl中不带参数的split等价于split /\s+/, $_, 0,会自动忽略开头的空白,并且丢弃末尾的空字符串,行为和Python的split()完全一致:
#!/usr/bin/perl -w use strict; my $lines=0; my $words=0; my $bytes=0; while (<>) { $lines++; $bytes += length; $words += split; # 默认split,自动忽略首尾空白、合并连续空白 } print "$lines $words $bytes\n";
方案2:手动去除首尾空白后分割
如果需要显式控制,可以先去除行首尾的空白,再分割统计:
#!/usr/bin/perl -w use strict; my $lines=0; my $words=0; my $bytes=0; while (<>) { $lines++; $bytes += length; s/^\s+|\s+$//g; # 移除首尾所有空白 $words += split /\s+/ if length; # 仅当处理后非空时统计单词数 } print "$lines $words $bytes\n";
修复后运行,Perl版本的统计结果会和标准wc、Python版本完全一致。
内容的提问来源于stack exchange,提问作者Paul Schutte
相关产品推荐
相关产品推荐

