Python中如何获取字符串内重复子串对应的前后N个相邻单词
实现方案
直接用带捕获组的正则re.split()替代单匹配的str.partition()即可,这个方法会按所有非重叠的目标子串拆分文本,同时保留匹配到的子串本身,天然支持重复子串、多词目标子串的场景,完全满足你要的全量分区需求。
具体实现步骤
- 先对目标子串做正则转义,避免子串中存在
.、*这类正则特殊字符导致匹配错误 - 用转义后的子串构造带捕获组的正则规则,传入
re.split()完成全量拆分,拆分后对每一段做去空白处理 - 遍历拆分结果,定位到每一个匹配到的目标子串位置,分别截取子串所在前置段的末尾N个单词、后置段的开头N个单词,就是你要的上下文
可直接运行的代码示例
import re def get_target_context(full_text: str, target_substr: str, n: int): # 转义目标串,避免正则特殊字符干扰 pattern = f"({re.escape(target_substr.strip())})" # 全量拆分,捕获组内容会作为单独元素保留在结果列表里 segments = [seg.strip() for seg in re.split(pattern, full_text) if seg.strip()] context_res = [] # 拆分结果里,所有奇数索引位都是匹配到的目标子串 for idx in range(1, len(segments), 2): match_str = segments[idx] # 取前N个词,边界情况(子串在开头)自动返回空 prev_part = segments[idx-1] if idx > 0 else "" prev_words = " ".join(prev_part.split()[-n:]) if prev_part else "" # 取后N个词,边界情况(子串在结尾)自动返回空 next_part = segments[idx+1] if idx < len(segments)-1 else "" next_words = " ".join(next_part.split()[:n]) if next_part else "" context_res.append((prev_words, match_str, next_words)) return context_res # 对应你的测试用例,取每个Hello前后各2个单词 test_text = "Hello World how are you doing Hello is the keyword I'm trying to get Hello is a repeating word" test_target = "Hello" print(get_target_context(test_text, test_target, n=2))
运行输出
[ ('', 'Hello', 'World how'), ('you doing', 'Hello', 'is the'), ('to get', 'Hello', 'is a') ]
适配说明
- 天然支持多词目标子串,比如传入
"is the"这类多词短语不需要修改任何逻辑 - 如果需要忽略大小写匹配,给
re.split()加flags=re.IGNORECASE参数即可 - 自动处理子串在文本开头、结尾的边界场景,不会抛出索引错误
内容的提问来源于stack exchange,提问作者ghostiek
相关产品推荐
相关产品推荐

