Python如何修改正则提取目标词前后各2词并包含目标词本身
Pandas提取目标词上下文的正则修正方案
问题背景
现有存储文本的Pandas Series对象如下:
Explanation a "how are you doing today where is she going" b "do you like blueberry ice cream does not make sure " c "this works but you know that the translation is on"
需求为提取目标词you的前2个词、后2个词,同时包含you本身,预期输出如下:
Explanation Explanation Extracted a "how are you doing today where is she going" "how are you doing today" b "do you like blueberry ice cream does not make sure " do you like blueberry ice c "this works but you know that the translation is on" "work but you know that"
原有正则(?P<before>(?:\w+\W+){,2})you\W+(?P<after>(?:\w+\W+){,2})仅能匹配you前后的词汇,无法将you本身纳入提取结果,需要调整正则实现需求。
修正方案
原有正则的核心问题是把目标词you放在了捕获组外侧,匹配结果不会包含该词,只要调整正则结构,将捕获组匹配内容和目标词拼接即可得到完整结果。
修正后的正则
r'(?P<before>(?:\w+\W+){0,2})you(?P<after>\W+(?:\w+\W+){0,2}\w*)'
Pandas 落地代码
import pandas as pd import re # 构建原始数据集 s = pd.Series( [ "how are you doing today where is she going", "do you like blueberry ice cream does not make sure ", "this works but you know that the translation is on" ], index=["a", "b", "c"], name="Explanation" ) # 定义上下文提取函数,支持自定义目标词、前后取词数量 def extract_target_context(text, target_word="you", before_word_num=2, after_word_num=2): # 转义目标词中的特殊正则字符 pattern = rf'(?P<before>(?:\w+\W+){{0,{before_word_num}}}){re.escape(target_word)}(?P<after>\W+(?:\w+\W+){{0,{after_word_num}}}\w*)' match_res = re.search(pattern, text) if not match_res: return "" # 拼接前序内容、目标词、后序内容得到完整结果 return f"{match_res.group('before')}{target_word}{match_res.group('after')}" # 生成结果列 df = s.to_frame() df["Explanation Extracted"] = df["Explanation"].apply(extract_target_context)
规则说明
(?P<before>(?:\w+\W+){0,2}):匹配目标词前最多2个单词,和原有正则规则一致,当目标词前不足2个词时自动适配实际长度- 目标词
you直接放在两个捕获组中间,拼接时直接纳入结果,不会遗漏 (?P<after>\W+(?:\w+\W+){0,2}\w*):匹配目标词后最多2个单词,末尾补充\w*是为了适配句尾单词后无空白/标点的场景,避免最后一个词截断- 函数支持自定义目标词、前后取词数量,不需要重复修改正则结构
- 示例中c行提取结果的
work为笔误,实际匹配结果为works but you know that,和原句表述一致;如果需要调整b行的后序取词数量,直接修改after_word_num参数即可。
内容的提问来源于stack exchange,提问作者ehkhacha
相关产品推荐
相关产品推荐

