You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何Bash程序运行飞快而Perl程序却如此缓慢?

为什么Bash实现比Perl更快?

问题重现

Bash实现代码

dict_fn='/usr/share/dict/american-english'
benchmark_repetitions=100
for i in `seq 1 $benchmark_repetitions` ; do
    readarray -t syllables < 'syllables.txt'
    for syllable in "${syllables[@]}" ; do
        matches=$(grep "$syllable" $dict_fn | wc -l)
        #echo "${syllable}: $matches"
    done
done

运行耗时:

$ time t.sh 

real    0m3,488s
user    0m4,103s
sys 0m0,910s

输入文件syllables.txt内容

art
bur
cic
dun
eik
fum
gaw
hiw
irc
jaw
kry
lus
mac
noq
old
pew
qla
rub
sur
tif
uks

Perl实现代码

use feature qw(say);
use strict;
use warnings;
use experimental qw(declared_refs refaliasing signatures);

{
    my $benchmark_repetitions = 100;
    for (1..$benchmark_repetitions) {
        my \@lines = read_dict();
        my \@syllables = read_syllables();
        for my $syllable (@syllables) {
            my @matches = grep /\Q$syllable/, @lines;
            my $num_matches = scalar @matches;
            #say "${syllable}: ${num_matches}";
        }
    }
}

sub read_syllables() {
    my $fn = 'syllables.txt';
    return read_file($fn);
}

sub read_dict() {
    my $fn = '/usr/share/dict/american-english'; # A file with more than 100,000 words
    return read_file($fn);
}

sub read_file( $fn ) {
    open ( my $fh, '<', $fn ) or die "Could not open file '$fn': $!";
    chomp(my @lines = <$fh>);
    close $fh;
    return \@lines;
}

运行耗时:

$ time p.pl

real    0m13,392s
user    0m13,369s
sys 0m0,021s

原因分析

Perl代码的核心低效点

  1. 重复加载大文件:脚本在100次循环中,每次都重新读取10万行的字典文件并构建数组,这带来了大量不必要的内存分配和IO重复开销——哪怕有系统缓存,重复构建数组的过程也会消耗大量时间。
  2. 低效的匹配方式:对每个音节,用Perl内置grep结合正则遍历整个数组。Perl的正则匹配是解释型逻辑,逐行遍历数组的操作效率远低于专门优化的文本搜索工具。

Bash代码的优势

  1. 利用系统文件缓存:第一次调用grep读取字典文件后,操作系统会将文件内容缓存到内存,后续所有grep调用都直接从内存读取,几乎无磁盘IO开销。
  2. GNU grep的高度优化:GNU grep是C语言实现的专业文本搜索工具,内置Boyer-Moore等高效匹配算法,经过几十年性能打磨,子串匹配效率远高于Perl的解释型数组遍历。

优化后的Perl代码

解决核心问题:仅加载一次字典和音节,改用更高效的子串匹配方式(index函数比正则更快):

use feature qw(say);
use strict;
use warnings;
use experimental qw(declared_refs refaliasing signatures);
use List::Util qw(sum0); # 引入sum0提升计数效率

# 仅加载一次字典和音节,避免重复IO
my \@dict_lines = read_dict();
my \@syllables = read_syllables();

my $benchmark_repetitions = 100;
for (1..$benchmark_repetitions) {
    for my $syllable (@syllables) {
        # 用index做子串匹配,比正则grep更高效
        my $num_matches = sum0 map { index($_, $syllable) != -1 ? 1 : 0 } @dict_lines;
        #say "${syllable}: ${num_matches}";
    }
}

sub read_syllables() {
    my $fn = 'syllables.txt';
    return read_file($fn);
}

sub read_dict() {
    my $fn = '/usr/share/dict/american-english';
    return read_file($fn);
}

sub read_file( $fn ) {
    open ( my $fh, '<', $fn ) or die "Could not open file '$fn': $!";
    chomp(my @lines = <$fh>);
    close $fh;
    return \@lines;
}

优化后,Perl脚本的运行速度会显著提升,甚至超过Bash版本——既避免了重复加载文件的开销,又用更高效的匹配逻辑替代了原有的低效遍历。

内容的提问来源于stack exchange,提问作者Håkon Hægland

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 04:30:30