为何Perl的File::Map比File::Slurp性能差这么多?
问题:File::Map处理大文件时性能暴跌,是用法问题还是模块本身缺陷?
我原本想用mmap搜索多GB级文件,避免内存耗尽。但测试一个能完全放入内存的文件时,File::Slurp版本耗时不到一分钟,File::Map版本跑了数分钟还没完成,只能终止。
测试小文件后发现,File::Map的性能随文件增大持续下降(文件大小翻倍,耗时变为4倍),而File::Slurp性能相对稳定(文件大小翻倍,耗时也翻倍)。
是我对该模块用法有误,还是File::Map处理大文件时本就很慢?
for n in 1 4 16 32 64 128 256 512 4096; do seq $n | xargs -I@ seq 100000 > data ls -l data time perl -MFile::Slurp -e ' $s = read_file("data"); $re = qr/^(99999|12345|4325|11111|50000)$/m; while ($s =~ m/$re/g){ ++$matches } print $matches; ' time perl -MFile::Map=:all -e ' map_file $s, "data"; advise $s, "sequential"; $re = qr/^(99999|12345|4325|11111|50000)$/m; while ($s =~ m/$re/g){ ++$matches } print $matches; ' done
| n | 文件大小 | 匹配次数 | 用户耗时(Slurp) | 用户耗时(Slurp)/n | 系统耗时(Slurp) | 系统耗时(Slurp)/n | 用户耗时(Map) | 用户耗时(Map)/n | 系统耗时(Map) | 系统耗时(Map)/n |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 588895 | 5 | 0.033 | 0.033 | 0.007 | 0.007 | 0.014 | 0.014 | 0.001 | 0.001 |
| 4 | 2355580 | 20 | 0.051 | 0.013 | 0.007 | 0.002 | 0.032 | 0.008 | 0.005 | 0.001 |
| 16 | 9422320 | 80 | 0.109 | 0.007 | 0.015 | 0.001 | 0.138 | 0.009 | 0.012 | 0.001 |
| 32 | 18844640 | 160 | 0.184 | 0.005 | 0.024 | 0.001 | 0.400 | 0.013 | 0.021 | 0.001 |
| 64 | 37689280 | 320 | 0.328 | 0.005 | 0.049 | 0.001 | 2.666 | 0.042 | 4.305 | 0.067 |
| 128 | 75378560 | 640 | 0.629 | 0.005 | 0.079 | 0.001 | 10.014 | 0.078 | 17.638 | 0.138 |
| 256 | 150757120 | 1280 | 1.220 | 0.005 | 0.162 | 0.001 | 40.237 | 0.157 | 73.829 | 0.288 |
| 512 | 301514240 | 2560 | 2.423 | 0.005 | 0.323 | 0.001 | 158.729 | 0.310 | 302.041 | 0.590 |
| 4096 | 2412113920 | 20480 | 19.468 | 0.005 | 2.424 | 0.001 | ? | ? | ? | ? |
按照@TLP的建议,不再手动从ls和time输出计算表格,以下是Perl Benchmark版本的测试结果(已省略警告输出),同样显示File::Slurp性能不受文件大小影响,而File::Map性能随文件增大下降:
#!/bin/bash for n in 1 4 16 32 64 128 256 512; do seq $n | xargs -I@ seq 100000 > data$n done perl -MBenchmark=cmpthese -MFile::Slurp -MFile::Map=:all -e ' @n = (1,4,16,32,64,128,256,512); sub test_slurp { my ($s,$re,$matches); $s = read_file($f); $re = qr/^(99999|12345|4325|11111|50000)$/m; while ($s =~ m/$re/g){ ++$matches } } sub test_map { my ($mm,$re,$matches); map_file $mm, $f; advise $mm, "sequential"; $re = qr/^(99999|12345|4325|11111|50000)$/m; while ($mm =~ m/$re/g){ ++$matches } } for $n (@n) { $f = "data$n"; cmpthese(-1, { "map($n)" => \&test_map, "slurp($n)" => \&test_slurp }); } '
速率 map(1) slurp(1) map(1) 198次/秒 -- -1% slurp(1) 200次/秒 1% -- 速率 map(4) slurp(4) map(4) 38.3次/秒 -- -20% slurp(4) 48.1次/秒 26% -- 速率 map(16) slurp(16) map(16) 6.60次/秒 -- -48% slurp(16) 12.6次/秒 91% -- 速率 map(32) slurp(32) map(32) 1.98次/秒 -- -62% slurp(32) 5.17次/秒 161% -- 耗时/迭代 map(64) slurp(64) map(64) 7.93秒 -- -96% slurp(64) 0.350秒 2166% -- 耗时/迭代 map(128) slurp(128) map(128) 31.6秒 -- -98% slurp(128) 0.730秒 4233% -- 耗时/迭代 map(256) slurp(256) map(256) 129秒 -- -99% slurp(256) 1.55秒 8244% -- 耗时/迭代 map(512) slurp(512) map(512) 521秒 -- -99% slurp(512) 2.82秒 18372% --
内容的提问来源于stack exchange,提问作者jhnc
相关产品推荐
相关产品推荐

