如何调整正则表达式,在前后文本可变时提取中间目标内容?
问题描述
我需要从一段长文本中提取目标名称(示例为"Melinda Gates"),已知目标内容邻近的固定核心文本片段:
- 目标前的邻近固定片段:
"came home last night to" - 目标后的邻近固定片段:
"his wife of several years"
长文本示例:
Before he came home last night to Melinda Gates, his wife of several years who loves him dearly
当前使用的正则模式为:
import re before_text = "He came home last night to" after_text = "his wife of several years" pattern = fr"{re.escape(before_text)}(.*?){re.escape(after_text)}"
但当before_text或after_text前后出现额外内容(比如before_text变为"While he was drunk and came home last night to",after_text变为"his wife of several years who grew up in Poland")时,正则会失效——因为当前正则匹配的是完整的before/after文本,而非邻近目标的固定核心片段。
解决方案
核心思路:只匹配目标内容前后的固定核心片段,忽略核心片段之外的所有可变内容。
调整后的正则实现
直接基于固定核心片段构建正则,用通配符匹配核心片段前后的任意内容:
import re # 目标前后的固定核心片段 fixed_prefix = "came home last night to" fixed_suffix = "his wife of several years" # 构建正则:匹配任意内容 + 固定前缀 + 目标内容 + 固定后缀 + 任意内容 pattern = fr".*{re.escape(fixed_prefix)}(.*?){re.escape(fixed_suffix)}.*" long_text = "Before he came home last night to Melinda Gates, his wife of several years who loves him dearly" match = re.search(pattern, long_text) if match: target_name = match.group(1).strip() # 去除目标内容前后的空格、标点 print(target_name) # 输出: Melinda Gates
关键细节
.*用于匹配固定核心片段前后的任意可变内容,不管前后新增多少文字都不影响匹配re.escape()处理固定核心片段,避免其中的空格、普通字符被正则误解析为特殊语法.strip()用于清理目标内容前后可能附带的空格、逗号等无关符号
进阶优化(可选)
如果需要确保固定核心片段是完整短语(避免被部分匹配),可以添加单词边界\b:
pattern = fr".*\b{re.escape(fixed_prefix)}\b(.*?)\b{re.escape(fixed_suffix)}\b.*"
这样能防止类似"came home last night toXYZ"这类不符合预期的部分匹配。
内容的提问来源于stack exchange,提问作者minuscler
相关产品推荐
相关产品推荐

