Python正则提取两关键词间文本如何匹配到首个结束关键词
使用Python re模块开展正则文本提取时,目标是提取字符串中起始关键词COMPUTATION OF DAMAGES与结束关键词DIMOPOULOS INJURY之间的文本,初始实现代码如下:
import re a = """COMPUTATION OF DAMAGES Plaintiff Maurice’s computation of damages to date includes all the above related medical specials, totaling $98,429.00. The Minimally Invasive Hand Institute $49,949.00 Interventional Pain & Spine Institute $1,190.00 Premier Physical Therapy $8,600.00 Clinical Neurology Specialist $3,090.00 Red Rock Surgery Center $34,510.00 DIMOPOULOS INJURY This is the bill of 1 2 3 4 5 6 7 8 DIMOPOULOS INJURY """ word1 = "COMPUTATION OF DAMAGES" word2 = "DIMOPOULOS INJURY" result = re.search(word1 + '(.*)' + word2, a) print(result.group(1))
初始代码的预期输出为:
Plaintiff Maurice’s computation of damages to date includes all the above related medical specials, totaling $98,429.00. The Minimally Invasive Hand Institute $49,949.00 Interventional Pain & Spine Institute $1,190.00 Premier Physical Therapy $8,600.00 Clinical Neurology Specialist $3,090.00 Red Rock Surgery Center $34,510.00
由于Python正则默认采用贪婪匹配模式,(.*)会尽可能匹配最长的符合规则的内容,最终会匹配到字符串中最后一次出现的DIMOPOULOS INJURY位置,无法得到预期结果。
方案1:修改正则为非贪婪(惰性)匹配
这是改动量最小的修复方式:在通配量词*后追加?,将(.*)修改为(.*?),即可将匹配逻辑切换为非贪婪模式,匹配时会尽可能少匹配字符,遇到第一个符合结束关键词的位置就停止匹配。
如果待提取的两个关键词之间存在换行内容,需要额外添加re.DOTALL(别名re.S)匹配标识,让正则中的.可以匹配换行符。
注意:如果关键词本身包含正则特殊字符(如
.、*、(等),建议使用re.escape()对关键词做转义处理,避免正则解析逻辑出错。
修改后的可运行代码:
import re a = """COMPUTATION OF DAMAGES Plaintiff Maurice’s computation of damages to date includes all the above related medical specials, totaling $98,429.00. The Minimally Invasive Hand Institute $49,949.00 Interventional Pain & Spine Institute $1,190.00 Premier Physical Therapy $8,600.00 Clinical Neurology Specialist $3,090.00 Red Rock Surgery Center $34,510.00 DIMOPOULOS INJURY This is the bill of 1 2 3 4 5 6 7 8 DIMOPOULOS INJURY """ word1 = "COMPUTATION OF DAMAGES" word2 = "DIMOPOULOS INJURY" result = re.search(re.escape(word1) + r'(.*?)' + re.escape(word2), a, flags=re.DOTALL) print(result.group(1).strip())
方案2:使用原生字符串切片实现(无正则兼容问题)
如果不想处理正则匹配模式的兼容问题,可以直接用Python字符串内置的查找方法实现截取,逻辑更直观,完全不存在贪婪/非贪婪的匹配歧义,执行效率也更高。
核心逻辑:
- 先查找起始关键词的结束位置,作为切片起点
- 从切片起点往后查找第一个结束关键词的起始位置,作为切片终点
- 直接切出两个位置之间的内容即可
实现代码:
a = """COMPUTATION OF DAMAGES Plaintiff Maurice’s computation of damages to date includes all the above related medical specials, totaling $98,429.00. The Minimally Invasive Hand Institute $49,949.00 Interventional Pain & Spine Institute $1,190.00 Premier Physical Therapy $8,600.00 Clinical Neurology Specialist $3,090.00 Red Rock Surgery Center $34,510.00 DIMOPOULOS INJURY This is the bill of 1 2 3 4 5 6 7 8 DIMOPOULOS INJURY """ word1 = "COMPUTATION OF DAMAGES" word2 = "DIMOPOULOS INJURY" start_idx = a.find(word1) + len(word1) end_idx = a.find(word2, start_idx) result = a[start_idx:end_idx].strip() print(result)
该方案不需要依赖正则模块,对包含特殊字符的关键词也不需要额外转义,适配绝大多数固定关键词截取的场景。
内容的提问来源于stack exchange,提问作者Manu Raj

