You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则提取两关键词间文本如何匹配到首个结束关键词

问题场景

使用Python re模块开展正则文本提取时,目标是提取字符串中起始关键词COMPUTATION OF DAMAGES与结束关键词DIMOPOULOS INJURY之间的文本,初始实现代码如下:

import re
a = """COMPUTATION OF DAMAGES Plaintiff Maurice’s computation of damages to date includes all the above related medical specials, totaling $98,429.00. The Minimally Invasive Hand Institute $49,949.00 Interventional Pain & Spine Institute $1,190.00 Premier Physical Therapy $8,600.00 Clinical Neurology Specialist $3,090.00 Red Rock Surgery Center $34,510.00 DIMOPOULOS INJURY This is the bill of 1 2 3 4 5 6 7  8 DIMOPOULOS INJURY """
word1 = "COMPUTATION OF DAMAGES"
word2 = "DIMOPOULOS INJURY"
result = re.search(word1 + '(.*)' + word2, a)
print(result.group(1))

初始代码的预期输出为:

Plaintiff Maurice’s computation of damages to date includes all the above related medical specials, totaling $98,429.00. The Minimally Invasive Hand Institute $49,949.00 Interventional Pain & Spine Institute $1,190.00 Premier Physical Therapy $8,600.00 Clinical Neurology Specialist $3,090.00 Red Rock Surgery Center $34,510.00

由于Python正则默认采用贪婪匹配模式,(.*)会尽可能匹配最长的符合规则的内容,最终会匹配到字符串中最后一次出现的DIMOPOULOS INJURY位置,无法得到预期结果。

可行解决方案

方案1:修改正则为非贪婪(惰性)匹配

这是改动量最小的修复方式:在通配量词*后追加?,将(.*)修改为(.*?),即可将匹配逻辑切换为非贪婪模式,匹配时会尽可能少匹配字符,遇到第一个符合结束关键词的位置就停止匹配。
如果待提取的两个关键词之间存在换行内容,需要额外添加re.DOTALL(别名re.S)匹配标识,让正则中的.可以匹配换行符。

注意:如果关键词本身包含正则特殊字符(如.、*、(等),建议使用re.escape()对关键词做转义处理,避免正则解析逻辑出错。

修改后的可运行代码:

import re
a = """COMPUTATION OF DAMAGES Plaintiff Maurice’s computation of damages to date includes all the above related medical specials, totaling $98,429.00. The Minimally Invasive Hand Institute $49,949.00 Interventional Pain & Spine Institute $1,190.00 Premier Physical Therapy $8,600.00 Clinical Neurology Specialist $3,090.00 Red Rock Surgery Center $34,510.00 DIMOPOULOS INJURY This is the bill of 1 2 3 4 5 6 7  8 DIMOPOULOS INJURY """
word1 = "COMPUTATION OF DAMAGES"
word2 = "DIMOPOULOS INJURY"
result = re.search(re.escape(word1) + r'(.*?)' + re.escape(word2), a, flags=re.DOTALL)
print(result.group(1).strip())

方案2:使用原生字符串切片实现(无正则兼容问题)

如果不想处理正则匹配模式的兼容问题,可以直接用Python字符串内置的查找方法实现截取,逻辑更直观,完全不存在贪婪/非贪婪的匹配歧义,执行效率也更高。
核心逻辑:

  • 先查找起始关键词的结束位置,作为切片起点
  • 从切片起点往后查找第一个结束关键词的起始位置,作为切片终点
  • 直接切出两个位置之间的内容即可

实现代码:

a = """COMPUTATION OF DAMAGES Plaintiff Maurice’s computation of damages to date includes all the above related medical specials, totaling $98,429.00. The Minimally Invasive Hand Institute $49,949.00 Interventional Pain & Spine Institute $1,190.00 Premier Physical Therapy $8,600.00 Clinical Neurology Specialist $3,090.00 Red Rock Surgery Center $34,510.00 DIMOPOULOS INJURY This is the bill of 1 2 3 4 5 6 7  8 DIMOPOULOS INJURY """
word1 = "COMPUTATION OF DAMAGES"
word2 = "DIMOPOULOS INJURY"

start_idx = a.find(word1) + len(word1)
end_idx = a.find(word2, start_idx)
result = a[start_idx:end_idx].strip()
print(result)

该方案不需要依赖正则模块,对包含特殊字符的关键词也不需要额外转义,适配绝大多数固定关键词截取的场景。

内容的提问来源于stack exchange,提问作者Manu Raj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 15:06:23