Perl数组匹配文件内容异常:仅返回一条结果的问题排查
问题排查与修正
需求描述
实现File1.txt与File2.txt的内容匹配:仅对比File1中>WORDS>关键字后的语句与File2中的内容,匹配成功的条目需输出File1中的完整对应行,预期输出如下:
Other >WORDS> Other than that?. Lunch >WORDS> when do you want to have lunch?.
文件内容
- File1.txt内容:
----------------------------------------------- Words | Keyword | Sentence ----------------------------------------------- Lunch >WORDS> when do you want to have lunch?. Hate >WORDS> I hate you. Other >WORDS> Other than that?
- File2.txt内容:
I love you. Other than that?. I like you. when do you want to have lunch?.
当前问题
编写的Perl代码仅返回一条匹配结果:
Other >WORDS> Other than that?
原代码实现
use strict; use warnings; use diagnostics; use Data::Dumper; use 5.010; my $new= 'File1.txt'; #read File1 my $old= 'File2.txt'; #read File2 my $string1; my $string2; my @new_array; my @old_array; my $string11; my @array1; #--------------------------------------------------------------- # Main #--------------------------------------------------------------- open(NEW_FILE, "<", $new) || die "Cannot open file $new to Read! - $!"; open(OLD_FILE, "<", $old) || die "Cannot open file $old to Read! - $!"; while (<NEW_FILE>) { my $string1= $_; my $string11= $_; if ($string1=~ m/WORDS/){ #matching the Keyword >WORDS> $string1 = $'; #string1 will take after >WORDS> $string11 = $_; #string11 will take the full. push (@new_array, ($string1)); #string1 = @new_array push (@array1, ($string11)); }} #string11 = @array1 while (<OLD_FILE>) { my $string2= $_; if ($string2 =~ m/WORDS/){ #matching the Keyword >WORDS> $string2 = $'; #string2 will take after >WORDS> push (@old_array, ($string2)); #string2 = @old_array }} #------Do comparison between new file and old file. (only after WORDS) my @intersection =(); my @unintersection = (); my %hash1 = map{$_ => 1} @old_array; foreach (@new_array){ if (defined $hash1{$_}){ push @intersection, $_; #this one will take the same array between new and old } else { push @unintersection, $_; #this one will take the new array only. So, will read this one. }}
截至此处,打印@unintersection可得到:
Other than that? when do you want to have lunch?.
后续对比代码:
my @same(); my @not_same= (); my %hash2 = map{$_ => 1} @unintersection; foreach (@array1) { if (@array1 = m/WORDS/){ @array1 = $'; if (defined $hash2{$_}) { @array1 = $_; push @same, $_; } else { push @not_same, $_;}}} print @same; print @not_same; close(NEW_FILE); close(OLD_FILE); close(NEW_OUTPUT_FILE);
核心问题点
- File2处理逻辑错误:原代码错误判断File2行是否包含
WORDS,但File2中根本无该关键字,导致@old_array为空,哈希匹配完全失效。 - 循环破坏原数组:后续对比代码中,
@array1 = m/WORDS/和@array1 = $'直接替换原数组,导致循环仅执行一次就丢失剩余数据。 - 未处理字符串差异:File1与File2的对应句子存在换行符、末尾标点/空格的细微差异,直接全字符匹配会失败。
修正后的代码
use strict; use warnings; use 5.010; my $file1 = 'File1.txt'; my $file2 = 'File2.txt'; # 读取File2内容,预处理后存入哈希 my %file2_sentences; open my $fh2, '<', $file2 or die "无法打开$file2: $!"; while (<$fh2>) { chomp; # 移除换行符 s/^\s+|\s+$//g; # 去除首尾空白 next if length == 0; # 跳过空行 $file2_sentences{lc($_)} = 1; # 转小写存入,避免大小写匹配问题 } close $fh2; # 读取File1,匹配并输出结果 open my $fh1, '<', $file1 or die "无法打开$file1: $!"; while (<$fh1>) { next unless />(WORDS)>/; # 仅处理包含>WORDS>的行 my $full_line = $_; my $sentence = $'; # 提取>WORDS>之后的内容 chomp $sentence; $sentence =~ s/^\s+|\s+$//g; # 预处理句子空白 # 处理File1与File2的标点差异,双向匹配 if (exists $file2_sentences{lc($sentence)} || exists $file2_sentences{lc("$sentence.")}) { print $full_line; } } close $fh1;
代码说明
- File2预处理:统一处理换行符、首尾空白,转小写存入哈希,实现快速查找。
- File1逐行处理:仅筛选目标行,提取关键字后内容并预处理,同时兼容句子末尾的标点差异。
- 避免数据破坏:不修改原数组,循环逐行遍历所有目标数据。
运行结果
执行后将输出预期的两条匹配结果:
Lunch >WORDS> when do you want to have lunch?. Other >WORDS> Other than that?
内容的提问来源于stack exchange,提问作者DM 256
相关产品推荐
相关产品推荐

