You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何修改正则提取目标词前后各2词并包含目标词本身

Pandas提取目标词上下文的正则修正方案

问题背景

现有存储文本的Pandas Series对象如下:

Explanation 
a      "how are you doing today where is she going" 
b      "do you like blueberry ice cream does not make sure " 
c      "this works but you know that the translation is on"

需求为提取目标词you的前2个词、后2个词,同时包含you本身,预期输出如下:

Explanation                                                    Explanation Extracted
a      "how are you doing today where is she going"                  "how are you doing today"
b      "do you like blueberry ice cream does not make sure "         do you like blueberry ice 
c      "this works but you know that the translation is on"           "work but you know that"

原有正则(?P<before>(?:\w+\W+){,2})you\W+(?P<after>(?:\w+\W+){,2})仅能匹配you前后的词汇,无法将you本身纳入提取结果,需要调整正则实现需求。

修正方案

原有正则的核心问题是把目标词you放在了捕获组外侧,匹配结果不会包含该词,只要调整正则结构,将捕获组匹配内容和目标词拼接即可得到完整结果。

修正后的正则

r'(?P<before>(?:\w+\W+){0,2})you(?P<after>\W+(?:\w+\W+){0,2}\w*)'

Pandas 落地代码

import pandas as pd
import re

# 构建原始数据集
s = pd.Series(
    [
        "how are you doing today where is she going",
        "do you like blueberry ice cream does not make sure ",
        "this works but you know that the translation is on"
    ],
    index=["a", "b", "c"],
    name="Explanation"
)

# 定义上下文提取函数,支持自定义目标词、前后取词数量
def extract_target_context(text, target_word="you", before_word_num=2, after_word_num=2):
    # 转义目标词中的特殊正则字符
    pattern = rf'(?P<before>(?:\w+\W+){{0,{before_word_num}}}){re.escape(target_word)}(?P<after>\W+(?:\w+\W+){{0,{after_word_num}}}\w*)'
    match_res = re.search(pattern, text)
    if not match_res:
        return ""
    # 拼接前序内容、目标词、后序内容得到完整结果
    return f"{match_res.group('before')}{target_word}{match_res.group('after')}"

# 生成结果列
df = s.to_frame()
df["Explanation Extracted"] = df["Explanation"].apply(extract_target_context)

规则说明

  • (?P<before>(?:\w+\W+){0,2}):匹配目标词前最多2个单词,和原有正则规则一致,当目标词前不足2个词时自动适配实际长度
  • 目标词you直接放在两个捕获组中间,拼接时直接纳入结果,不会遗漏
  • (?P<after>\W+(?:\w+\W+){0,2}\w*):匹配目标词后最多2个单词,末尾补充\w*是为了适配句尾单词后无空白/标点的场景,避免最后一个词截断
  • 函数支持自定义目标词、前后取词数量,不需要重复修改正则结构
  • 示例中c行提取结果的work为笔误,实际匹配结果为works but you know that,和原句表述一致;如果需要调整b行的后序取词数量,直接修改after_word_num参数即可。

内容的提问来源于stack exchange,提问作者ehkhacha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 20:09:18