Python正则如何匹配最近右括号 删除以(see §/See §开头的括号内容
问题原因
你当前的正则匹配失效,本质是两个问题:
.*属于贪婪匹配,默认会匹配尽可能长的字符,所以会一路匹配到整个字符串最后一个)才停止,把中间所有不该删的内容都吞掉- 就算把
.*改成非贪婪写法.*?,也处理不了待删除片段里的嵌套括号(比如第一个要删的块里的(PN)),非贪婪匹配碰到第一个内部的)就会终止,没法完整抓全整个要删除的括号块。
可行解决方案
方案1:遍历计数法(无额外依赖,逻辑稳定)
不需要依赖复杂正则,直接遍历字符串,遇到(see §/(See §开头的括号块时,通过计数左右括号的层级,找到和开头括号配平的闭合右括号,直接跳过整段内容即可,任意深度的嵌套括号都能正确处理:
def remove_see_paragraphs(s: str) -> str: res = [] n = len(s) i = 0 while i < n: # 命中待删除括号块的开头 if s[i] == '(' and s[i+1:i+5] in ('see §', 'See §'): level = 1 i += 1 # 计数括号层级,找到匹配的闭合右括号 while i < n and level > 0: if s[i] == '(': level += 1 elif s[i] == ')': level -= 1 i += 1 # 跳过括号块后面多余的空格,避免出现连续空格 while i < n and s[i] == ' ': i += 1 else: res.append(s[i]) i += 1 return ''.join(res) # 测试验证 txt = '(i) Test text (see § 123.1 of this Proper Name (PN) subparagraph) and additional test text (see § 123.2 of this subparagraph) along with more test text (including a parenthetical) to test the text. (See § 125.3 of this subparagraph for even more test text.) (Hello World!)' print(remove_see_paragraphs(txt))
运行后输出和预期完全一致:(i) Test text and additional test text along with more test text (including a parenthetical) to test the text. (Hello World!)
方案2:递归正则法(需要安装第三方库)
Python标准库re不支持递归匹配,无法处理嵌套结构,可以安装第三方regex库,用递归规则匹配任意层级的括号块:
- 先安装依赖:
pip install regex - 实现代码:
import regex txt = '(i) Test text (see § 123.1 of this Proper Name (PN) subparagraph) and additional test text (see § 123.2 of this subparagraph) along with more test text (including a parenthetical) to test the text. (See § 125.3 of this subparagraph for even more test text.) (Hello World!)' # 递归正则匹配(see §开头、支持任意嵌套括号的完整块,同时匹配块后多余空格 pattern = r'\((?:s|S)ee §(?:[^()]|\((?:[^()]|(?R))*\))*\)\s*' res = regex.sub(pattern, '', txt) print(res)
注意事项
- 不要硬用标准库
re写处理嵌套括号的正则,标准库不支持递归/平衡组特性,写出来的规则漏匹配、误匹配概率极高,后续维护成本很大 - 如果不想安装第三方依赖,优先选第一种遍历计数的方案,逻辑透明好调试,不管括号嵌套多少层都能稳定运行。
内容的提问来源于stack exchange,提问作者txmountaineer
相关产品推荐
相关产品推荐

