如何编写正则表达式移除包含图、表引用内容的括号?
正则表达式修正方案
你原来的正则匹配规则覆盖范围过窄,仅能匹配单字符+空格+数字+转义序列的组合,无法适配figure/table这类关键词、and连接的多引用场景,以及编号中的特殊短横字符,所以无法命中目标内容。
场景1:删除整个括号及内部的引用内容
如果你的需求是把整段括号包裹的引用直接删掉,使用以下正则即可:r'\((?:figures?|tables?)\s[\d–]+(?:\s+and\s+[\d–]+)?\)'
代码示例
import re text_list = [ "this is a figure ref (figure 7–77)", "this is multiple refs (figures 6–28 and 6–29)", "this is a table ref (table 6–1)" ] processed_text = [re.sub(r'\((?:figures?|tables?)\s[\d–]+(?:\s+and\s+[\d–]+)?\)', '', item) for item in text_list]
处理后输出
[ "this is a figure ref ", "this is multiple refs ", "this is a table ref " ]
场景2:仅移除括号、保留内部引用内容
如果只是要去掉包裹的括号,把里面的引用内容留在原文中,用捕获组提取内容替换即可:
processed_text = [re.sub(r'\(((?:figures?|tables?)\s[\d–]+(?:\s+and\s+[\d–]+)?)\)', r'\1', item) for item in text_list]
处理后输出
[ "this is a figure ref figure 7–77", "this is multiple refs figures 6–28 and 6–29", "this is a table ref table 6–1" ]
内容的提问来源于stack exchange,提问作者MotzWanted
相关产品推荐
相关产品推荐

