You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则匹配分号开头注释(排除未转义引号包裹场景)

正则匹配分号开头的注释(排除被未转义双引号包裹的情况)

需求:匹配以分号开头的行尾注释,但如果分号被未转义的双引号同时包裹两侧,则不匹配该分号(仅绿色块标注的注释需要被匹配)。

规则说明:

  • 双引号通过重复""转义,转义后的双引号不具备“包裹分号以禁用注释”的作用;
  • 不平衡的双引号视为转义后的,同样不生效。

现有正则^(?>(?:""[^""\n]*""|[^;""\n]+)*)""?[^"";\n]*(;.*)存在缺陷,无法正确处理最后一个测试向量行末尾的转义双引号。

测试向量(与示意图一致):

Peekaboo ; A comment starts with a semicolon and continues till the EOL
Unless the semicolon is surrounded by dquotes ”Don’t do it ; here” ;but match me; once
Im not surrounded ”so pay attention to me” ; ”peekaboo”
Im not surrounded ”so pay attention” to;me” ; ”peekaboo”
Im not surrounded ”so pay attention to me ; peekaboo
Dquote escapes a dquote so ”dont pay attention to ””me;here”” buster” do it ; here
Don’t pay attention to  ”””me;here””” but do ””it;here””
and ”dont do ””it;here”””  either ;peekaboo
but "pay attention to "it;here"" ;not here though
Simon said ”I like goats” then he added ”and sheep;” ;a good comment is ”here
Simon said ”I like goats” then he added ”and sheep;” dont do it here
Simon said ””I like goats;”peekaboo
Simon said ”I like goats;””peekaboo

修正后的正则

^(?>(?:""[^"\n]*""|[^;"\n]+)*)(?:""?[^;"\n]*)(;.*)

逻辑说明

原正则的""?[^"";\n]*部分在处理行尾连续转义双引号时会出错,修正后:

  1. 原子组(?>(?:""[^"\n]*""|[^;"\n]+)*)负责匹配所有非注释内容,包括转义的双引号对和普通文本;
  2. (?:""?[^;"\n]*)替换原逻辑,确保单个未闭合引号(视为转义)或转义对残留都能被正确跳过,精准定位到第一个有效的注释起始分号;
  3. 最终捕获(;.*)作为目标注释。

该正则可正确处理所有测试向量,包括最后一行的转义双引号场景。

内容的提问来源于stack exchange,提问作者Pavel Stepanek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 11:35:38