You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则如何匹配最近右括号 删除以(see §/See §开头的括号内容

问题原因

你当前的正则匹配失效,本质是两个问题:

  • .*属于贪婪匹配,默认会匹配尽可能长的字符,所以会一路匹配到整个字符串最后一个)才停止,把中间所有不该删的内容都吞掉
  • 就算把.*改成非贪婪写法.*?,也处理不了待删除片段里的嵌套括号(比如第一个要删的块里的(PN)),非贪婪匹配碰到第一个内部的)就会终止,没法完整抓全整个要删除的括号块。
可行解决方案

方案1:遍历计数法(无额外依赖,逻辑稳定)

不需要依赖复杂正则,直接遍历字符串,遇到(see §/(See §开头的括号块时,通过计数左右括号的层级,找到和开头括号配平的闭合右括号,直接跳过整段内容即可,任意深度的嵌套括号都能正确处理:

def remove_see_paragraphs(s: str) -> str:
    res = []
    n = len(s)
    i = 0
    while i < n:
        # 命中待删除括号块的开头
        if s[i] == '(' and s[i+1:i+5] in ('see §', 'See §'):
            level = 1
            i += 1
            # 计数括号层级,找到匹配的闭合右括号
            while i < n and level > 0:
                if s[i] == '(':
                    level += 1
                elif s[i] == ')':
                    level -= 1
                i += 1
            # 跳过括号块后面多余的空格,避免出现连续空格
            while i < n and s[i] == ' ':
                i += 1
        else:
            res.append(s[i])
            i += 1
    return ''.join(res)

# 测试验证
txt = '(i) Test text (see § 123.1 of this Proper Name (PN) subparagraph) and additional test text (see § 123.2 of this subparagraph) along with more test text (including a parenthetical) to test the text. (See § 125.3 of this subparagraph for even more test text.) (Hello World!)'
print(remove_see_paragraphs(txt))

运行后输出和预期完全一致:
(i) Test text and additional test text along with more test text (including a parenthetical) to test the text. (Hello World!)

方案2:递归正则法(需要安装第三方库)

Python标准库re不支持递归匹配,无法处理嵌套结构,可以安装第三方regex库,用递归规则匹配任意层级的括号块:

  1. 先安装依赖:pip install regex
  2. 实现代码:
import regex

txt = '(i) Test text (see § 123.1 of this Proper Name (PN) subparagraph) and additional test text (see § 123.2 of this subparagraph) along with more test text (including a parenthetical) to test the text. (See § 125.3 of this subparagraph for even more test text.) (Hello World!)'
# 递归正则匹配(see §开头、支持任意嵌套括号的完整块,同时匹配块后多余空格
pattern = r'\((?:s|S)ee §(?:[^()]|\((?:[^()]|(?R))*\))*\)\s*'
res = regex.sub(pattern, '', txt)
print(res)
注意事项
  • 不要硬用标准库re写处理嵌套括号的正则,标准库不支持递归/平衡组特性,写出来的规则漏匹配、误匹配概率极高,后续维护成本很大
  • 如果不想安装第三方依赖,优先选第一种遍历计数的方案,逻辑透明好调试,不管括号嵌套多少层都能稳定运行。

内容的提问来源于stack exchange,提问作者txmountaineer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 06:54:29