Perl处理.html文件时出现UTF-8编码错误如何解决
问题修复方案
你的编码错误由三个常见的Perl Unicode处理疏漏导致,按以下步骤修改即可解决:
问题根因
- 脚本未开启全局Unicode处理支持,字符串默认按字节流处理
readdir从操作系统读取的文件名是原始字节序列,没有解码为UTF-8字符串,带重音字符的文件名匹配、读写都会出现编码错位- 输出文件打开时机错误,未完成文件类型过滤就提前创建输出句柄,会生成大量无用空文件,也可能干扰编码写入逻辑
修复后的完整代码
use strict; use warnings; use utf8; # 声明脚本自身源码为UTF-8编码 use Encode qw(decode); use File::Spec::Functions qw/ splitdir rel2abs /; # 全局设置默认文件编码为UTF-8,标准输入输出也用UTF-8 use open qw(:std :encoding(UTF-8)); my ($inputfile, $outputfile, $dir); $dir = '.'; opendir(DIR, $dir); while (my $raw_filename = readdir(DIR)) { # 把原始字节文件名解码为UTF-8字符串 my $inputfile = decode('UTF-8', $raw_filename); # 先过滤要处理的文件,再进行后续操作 next unless (-f "$dir/$inputfile"); next unless ($inputfile =~ m/_not_centered\.html$/); # 生成输出文件名 $outputfile = $inputfile; $outputfile =~ s/_not_centered//; # 现在再打开输入输出文件 open(my $ifh, '<', $inputfile) or die "打开输入文件失败 $inputfile: $!"; open(my $ofh, '>', $outputfile) or die "打开输出文件失败 $outputfile: $!"; while(<$ifh>) { if(/(<h2)(.*?)/) { print $ofh "$1 style=\"text-align: center;\"$2"; }else{ print $ofh $_; } } close $ifh; close $ofh; } closedir(DIR);
关键修改说明
- 新增
use utf8声明,告知Perl脚本自身的源码是UTF-8编码 - 新增
use open qw(:std :encoding(UTF-8))全局声明,所有文件读写默认用UTF-8编码,无需重复在open语句里指定 - 对
readdir返回的原始文件名做UTF-8解码,转换为Perl可识别的Unicode字符串,避免文件名编码错误 - 调整操作顺序,先完成文件过滤再创建输出文件,避免生成无用空文件
- 新增文件打开的错误校验,方便排查后续可能出现的权限、文件不存在类问题
补充:如果你确认待处理的HTML文件实际是Latin1(ISO-8859-1)编码而非UTF-8,只需要把
use open行的编码改为encoding(ISO-8859-1)即可,输出仍会自动转为UTF-8格式。
内容的提问来源于stack exchange,提问作者AlMa
相关产品推荐
相关产品推荐

