You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何获取字符串内重复子串对应的前后N个相邻单词

实现方案

直接用带捕获组的正则re.split()替代单匹配的str.partition()即可,这个方法会按所有非重叠的目标子串拆分文本,同时保留匹配到的子串本身,天然支持重复子串、多词目标子串的场景,完全满足你要的全量分区需求。

具体实现步骤

  • 先对目标子串做正则转义,避免子串中存在.、*这类正则特殊字符导致匹配错误
  • 用转义后的子串构造带捕获组的正则规则,传入re.split()完成全量拆分,拆分后对每一段做去空白处理
  • 遍历拆分结果,定位到每一个匹配到的目标子串位置,分别截取子串所在前置段的末尾N个单词、后置段的开头N个单词,就是你要的上下文

可直接运行的代码示例

import re

def get_target_context(full_text: str, target_substr: str, n: int):
    # 转义目标串,避免正则特殊字符干扰
    pattern = f"({re.escape(target_substr.strip())})"
    # 全量拆分,捕获组内容会作为单独元素保留在结果列表里
    segments = [seg.strip() for seg in re.split(pattern, full_text) if seg.strip()]
    context_res = []

    # 拆分结果里,所有奇数索引位都是匹配到的目标子串
    for idx in range(1, len(segments), 2):
        match_str = segments[idx]
        # 取前N个词,边界情况(子串在开头)自动返回空
        prev_part = segments[idx-1] if idx > 0 else ""
        prev_words = " ".join(prev_part.split()[-n:]) if prev_part else ""
        # 取后N个词,边界情况(子串在结尾)自动返回空
        next_part = segments[idx+1] if idx < len(segments)-1 else ""
        next_words = " ".join(next_part.split()[:n]) if next_part else ""
        context_res.append((prev_words, match_str, next_words))
    
    return context_res

# 对应你的测试用例,取每个Hello前后各2个单词
test_text = "Hello World how are you doing Hello is the keyword I'm trying to get Hello is a repeating word"
test_target = "Hello"
print(get_target_context(test_text, test_target, n=2))

运行输出

[
    ('', 'Hello', 'World how'),
    ('you doing', 'Hello', 'is the'),
    ('to get', 'Hello', 'is a')
]

适配说明

  • 天然支持多词目标子串,比如传入"is the"这类多词短语不需要修改任何逻辑
  • 如果需要忽略大小写匹配,给re.split()加flags=re.IGNORECASE参数即可
  • 自动处理子串在文本开头、结尾的边界场景,不会抛出索引错误

内容的提问来源于stack exchange,提问作者ghostiek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 22:45:38