You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Perl如何用单个正则表达式将含%分隔符的行拆分为代码和注释

解决方案

我们可以通过可选非捕获组把边界符和注释部分包装为可选项,单个正则即可覆盖所有场景,不需要用|拼接两个分支,也不需要额外的%match哈希做默认值补全。


核心正则

/^(?<a1code>.*?)(?:(?<a2boundary>(?<!\\)%)(?<a3cmnt>.*))?$/

各部分逻辑说明:

  • ^(?<a1code>.*?):从行首开始匹配尽可能少的字符,到第一个未转义百分号之前为止,命名捕获为a1code
  • 外层(?: ... )?为非捕获可选组,表示内部的「边界符+注释」部分可以不存在
    • (?<a2boundary>(?<!\\)%):匹配未转义的百分号,命名捕获为a2boundary
    • (?<a3cmnt>.*):匹配百分号之后到行尾的所有内容,命名捕获为a3cmnt
  • 结尾的$确保匹配覆盖整行内容

修改后的完整代码

#!/usr/bin/env perl
use strict; use warnings;
print join('', 'perl ', $^V, "\n",);
use Data::Dumper qw(Dumper); $Data::Dumper::Sortkeys = 1;

my $count=0;
while(<DATA>)
{
    $count++;
    print "$count\t";
    chomp;
    print "|$_|\n";
    if($_=~/^(?<a1code>.*?)(?:(?<a2boundary>(?<!\\)%)(?<a3cmnt>.*))?$/)
    {
        # 未匹配到的捕获组直接赋值为空字符串即可
        my %result = (
            a1code => $+{a1code} // '',
            a2boundary => $+{a2boundary} // '',
            a3cmnt => $+{a3cmnt} // '',
        );
        print "from single regex match:\n";
        print Dumper \%result;
    }
    else
    {
        die "no match? coding error, should never get here";
    }
    print "------------------------------------------\n";
}

__DATA__
This is 100\% text and below you find an empty line.

abba 5\% %comment 9\% %Borgia
%all comment
%

运行结果和原实现完全一致,所有边界场景均能正确处理:空行、无未转义%的行、行首就是%的行、%后无内容的行都符合拆分要求。


内容的提问来源于stack exchange,提问作者Jacob Wegelin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 02:45:02