如何用正则匹配分号开头注释(排除未转义引号包裹场景)
正则匹配分号开头的注释(排除被未转义双引号包裹的情况)
需求:匹配以分号开头的行尾注释,但如果分号被未转义的双引号同时包裹两侧,则不匹配该分号(仅绿色块标注的注释需要被匹配)。
规则说明:
- 双引号通过重复
""转义,转义后的双引号不具备“包裹分号以禁用注释”的作用; - 不平衡的双引号视为转义后的,同样不生效。
现有正则^(?>(?:""[^""\n]*""|[^;""\n]+)*)""?[^"";\n]*(;.*)存在缺陷,无法正确处理最后一个测试向量行末尾的转义双引号。
测试向量(与示意图一致):
Peekaboo ; A comment starts with a semicolon and continues till the EOL Unless the semicolon is surrounded by dquotes ”Don’t do it ; here” ;but match me; once Im not surrounded ”so pay attention to me” ; ”peekaboo” Im not surrounded ”so pay attention” to;me” ; ”peekaboo” Im not surrounded ”so pay attention to me ; peekaboo Dquote escapes a dquote so ”dont pay attention to ””me;here”” buster” do it ; here Don’t pay attention to ”””me;here””” but do ””it;here”” and ”dont do ””it;here””” either ;peekaboo but "pay attention to "it;here"" ;not here though Simon said ”I like goats” then he added ”and sheep;” ;a good comment is ”here Simon said ”I like goats” then he added ”and sheep;” dont do it here Simon said ””I like goats;”peekaboo Simon said ”I like goats;””peekaboo
修正后的正则
^(?>(?:""[^"\n]*""|[^;"\n]+)*)(?:""?[^;"\n]*)(;.*)
逻辑说明
原正则的""?[^"";\n]*部分在处理行尾连续转义双引号时会出错,修正后:
- 原子组
(?>(?:""[^"\n]*""|[^;"\n]+)*)负责匹配所有非注释内容,包括转义的双引号对和普通文本; (?:""?[^;"\n]*)替换原逻辑,确保单个未闭合引号(视为转义)或转义对残留都能被正确跳过,精准定位到第一个有效的注释起始分号;- 最终捕获
(;.*)作为目标注释。
该正则可正确处理所有测试向量,包括最后一行的转义双引号场景。
内容的提问来源于stack exchange,提问作者Pavel Stepanek
相关产品推荐
相关产品推荐

