Perl中如何迭代遍历多行正则匹配块并解析为目标哈希结构
Perl 遍历多行正则匹配所有内容块实现方案
核心问题是原有代码在处理*multiline pattern(多行模式)*匹配时,块匹配逻辑为单次匹配,且使用贪婪匹配模式,仅能捕获单个内容块,无法遍历所有分段。
固定格式规则
- 每个内容块包含头部行,后接换行符,以及与头部长度相等的短横线分隔行
- 每个块固定以换行结尾,后接
(Number of results = \d+)格式行,再接换行符 - 普通键值对的等号前后固定为两个空格,匹配模式为
/ = / - 若行内容无key部分、直接以等号开头,则该值属于前一个key,需追加到对应数组中(如示例中
Alias字段) - 整个输入字符串固定以
--- END加换行符结尾
修正后可运行代码
#!/usr/bin/env perl use strict; use warnings; use Data::Dumper; local $/; my $output = <DATA>; my %hash; ($hash{'code'}, $hash{'description'}) = $output =~ /^RETCODE = (\d+)\s+(.*)\n/m; if ($hash{'code'} eq "0") { # 块匹配加/g全局修饰、非贪婪匹配.*?,循环遍历所有块 while ($output =~ /([^\n]+)\n-+\n(.*?)\n\n\(Number of results = (\d+)\)\n\n/msg) { my ($type, $data, $results) = ($1, $2, $3); $hash{$type}{results} = $results; my $previousKey = ""; # 逐行处理块内键值对 while ($data =~ /(.+)$/mg) { my $line = $1; $line =~ s/^ +//g; if ($line =~ /^\s*= /) { my ($value) = $line =~ /^\s*= (.*)$/; # 首次遇到多值key时,将原有标量转为数组 $hash{$type}{$previousKey} = [ $hash{$type}{$previousKey} ] unless ref($hash{$type}{$previousKey}); push (@{$hash{$type}{$previousKey}}, $value); } else { my ($key, $value) = split(/ = /, $line, 2); $hash{$type}{$key} = $value; $previousKey = $key; } } } print Dumper(\%hash); } __DATA__ +++ STAR-WARS 2020-01-01 00:00:00+00:00 S&W #00000000 %%SHOW NAME: Q=Kenobi;%% RETCODE = 0 Operation success In-universe information ----------------------- Species = Human Gender = Male television series of num = whatever (Number of results = 1) Personal Details ---------------- First Name = Obi-Wan Last Name = Kenobi Alias = Padawan = Jedi Knight = Jedi General = Jedi Master Points to other set of information = whatever (Number of results = 1) Other attribute --------------- Significant other = Satine Kryze Affiliation = Jedi Order = Galactic Republic = Rebel Alliance Occupation = Jedi (Number of results = 1) --- END
关键修改说明
- 块匹配正则添加
/g全局匹配修饰符,配合while循环时,正则引擎会自动记录上次匹配结束的位置,逐次向后查找所有符合规则的块,无需手动计算偏移量 - 将块内容匹配的贪婪模式
.*改为非贪婪模式.*?,避免贪婪匹配直接跨块吞掉中间所有内容,导致只匹配到最后一个块的问题 - 补全每个块的
results字段存储,和预期输出结构对齐 split键值对时增加第三个参数2,避免值本身包含=字符串时被意外拆分
运行后输出结构
$VAR1 = { 'In-universe information' => { 'Gender' => 'Male', 'Species' => 'Human', 'results' => '1', 'television series of num' => 'whatever' }, 'Other attribute' => { 'Affiliation' => [ 'Jedi Order', 'Galactic Republic', 'Rebel Alliance' ], 'Occupation' => 'Jedi', 'Significant other' => 'Satine Kryze', 'results' => '1' }, 'Personal Details' => { 'Alias' => [ 'Padawan', 'Jedi Knight', 'Jedi General', 'Jedi Master' ], 'First Name' => 'Obi-Wan', 'Last Name' => 'Kenobi', 'Points to other set of information' => 'whatever', 'results' => '1' }, 'code' => '0', 'description' => 'Operation success' };
内容的提问来源于stack exchange,提问作者h q
相关产品推荐
相关产品推荐

