You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改Python正则代码,仅移除((VERB)与)间的连续重复字符串?

问题:移除指定范围内的连续重复字符串

我编写了一段Python正则代码,意图仅移除位于((VERB)和)之间的连续重复字符串a nosotros,代码如下:

import re

input_text = "((VERB) saltar a nosotros a nosotros) a nosotros a nosotros a nosotros ((VERB)correr a nosotros) sdsdsd ((VERB) saltar a nosotros a nosotros)"

input_text = re.sub(r"\(\(VERB\)" + r"((?:\w\s*)+)" + r"\)", 
                    lambda x: re.sub(r"(a nosotros)\s*\1+", r"\1", x.group()), 
                    input_text)

print(input_text) # --> output

期望的输出结果为:

"((VERB) saltar a nosotros) a nosotros a nosotros a nosotros ((VERB)correr a nosotros) sdsdsd ((VERB) saltar a nosotros)"

当前代码未达到预期效果,请问需要如何修改?


修复方案

原代码的核心问题是外层正则未正确匹配((VERB) ... )的完整结构,同时内层替换逻辑可以进一步优化。修改后的代码如下:

import re

input_text = "((VERB) saltar a nosotros a nosotros) a nosotros a nosotros a nosotros ((VERB)correr a nosotros) sdsdsd ((VERB) saltar a nosotros a nosotros)"

input_text = re.sub(r"\(\(VERB\)\s*([^)]+)\)", 
                    lambda x: f"((VERB) {re.sub(r'(a nosotros)(?:\s+\1)+', r'\1', x.group(1).strip())})", 
                    input_text)

print(input_text)

关键修改点

  • 外层正则修正:将原正则改为\(\(VERB\)\s*([^)]+)\),精准匹配((VERB)后紧跟的),同时用\s*兼容((VERB)后有无空格的情况,([^)]+)可靠捕获目标范围内的内容。
  • 内层替换优化:用(a nosotros)(?:\s+\1)+匹配连续重复的a nosotros,非捕获组(?:...)避免多余捕获;通过x.group(1).strip()清除内容前后的冗余空格,再用f-string重新拼接成标准的((VERB) ... )格式。

运行修改后的代码,输出将与预期完全一致。


内容的提问来源于stack exchange,提问作者Matt095

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 05:53:19