You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何调整正则表达式,在前后文本可变时提取中间目标内容?

问题描述

我需要从一段长文本中提取目标名称(示例为"Melinda Gates"),已知目标内容邻近的固定核心文本片段:

  • 目标前的邻近固定片段:"came home last night to"
  • 目标后的邻近固定片段:"his wife of several years"

长文本示例:

Before he came home last night to Melinda Gates, his wife of several years who loves him dearly

当前使用的正则模式为:

import re

before_text = "He came home last night to"
after_text = "his wife of several years"
pattern = fr"{re.escape(before_text)}(.*?){re.escape(after_text)}"

但当before_text或after_text前后出现额外内容(比如before_text变为"While he was drunk and came home last night to",after_text变为"his wife of several years who grew up in Poland")时,正则会失效——因为当前正则匹配的是完整的before/after文本,而非邻近目标的固定核心片段。

解决方案

核心思路:只匹配目标内容前后的固定核心片段,忽略核心片段之外的所有可变内容。

调整后的正则实现

直接基于固定核心片段构建正则,用通配符匹配核心片段前后的任意内容:

import re

# 目标前后的固定核心片段
fixed_prefix = "came home last night to"
fixed_suffix = "his wife of several years"

# 构建正则:匹配任意内容 + 固定前缀 + 目标内容 + 固定后缀 + 任意内容
pattern = fr".*{re.escape(fixed_prefix)}(.*?){re.escape(fixed_suffix)}.*"

long_text = "Before he came home last night to Melinda Gates, his wife of several years who loves him dearly"
match = re.search(pattern, long_text)
if match:
    target_name = match.group(1).strip()  # 去除目标内容前后的空格、标点
    print(target_name)  # 输出: Melinda Gates

关键细节

  • .*用于匹配固定核心片段前后的任意可变内容,不管前后新增多少文字都不影响匹配
  • re.escape()处理固定核心片段,避免其中的空格、普通字符被正则误解析为特殊语法
  • .strip()用于清理目标内容前后可能附带的空格、逗号等无关符号

进阶优化(可选)

如果需要确保固定核心片段是完整短语(避免被部分匹配),可以添加单词边界\b:

pattern = fr".*\b{re.escape(fixed_prefix)}\b(.*?)\b{re.escape(fixed_suffix)}\b.*"

这样能防止类似"came home last night toXYZ"这类不符合预期的部分匹配。


内容的提问来源于stack exchange,提问作者minuscler

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 14:00:16