Python中从字符串索引获取完整单词及目标前后单词(规避正则)
不用正则获取匹配目标前后的单词
嘿,我完全理解你不想用正则处理大文件的想法——正则虽然灵活,但在处理海量文本时确实容易成为性能瓶颈。既然你已经通过正则拿到了每个匹配目标的起始索引,我们可以换用纯字符串操作和索引定位的方式来获取前后的单词,这样效率会高很多,而且也能满足你后续的其他需求。
核心思路
先预处理文本,把每个单词的起始索引、结束索引、单词内容都记录下来,形成一个有序的列表。之后针对每个匹配的起始索引,只需要在这个列表里快速定位到目标单词的位置,就能直接拿到前后的单词了。
具体实现代码
首先保留你原来的匹配逻辑,获取所有目标的起始索引:
import re text = 'hi this is john my name is john im bad boy' target = 'is john' # 你原有的匹配代码,获取每个匹配的起始索引 target_pattern = target.replace(' ', '[\s\n]*') target_re = re.compile(r'\b%s' % target_pattern, flags=re.I | re.X) indices = [m.start() for m in target_re.finditer(text)]
然后预处理文本,生成单词的位置映射:
# 生成每个单词的(起始索引, 结束索引, 单词内容)列表 word_positions = [] current_idx = 0 words = text.split() for word in words: # 找到当前单词在文本中的起始位置 start = text.find(word, current_idx) end = start + len(word) word_positions.append((start, end, word)) # 更新当前索引到单词结束后的下一个位置(跳过空格) current_idx = end + 1
接下来处理每个匹配索引,获取前后单词:
# 遍历每个匹配的起始索引 for match_start in indices: # 找到匹配目标所在的单词位置(这里match_start是"is"的起始索引) for i, (start, end, word) in enumerate(word_positions): if start <= match_start < end: # 前一个单词:当前单词的前一位 if i > 0: prev_word = word_positions[i-1][2] print(f"匹配'{target}'的前一个单词:{prev_word}") # 后一个单词:因为"is john"占2个单词,所以取当前位置+2的单词 if i + 2 < len(word_positions): next_word = word_positions[i+2][2] print(f"匹配'{target}'的后一个单词:{next_word}") break
性能优化:用二分查找快速定位
如果是处理超大文件,上面的遍历查找可以换成二分查找,进一步提升速度:
import bisect # 提取所有单词的起始索引,用于二分查找 start_indices = [pos[0] for pos in word_positions] for match_start in indices: # 用bisect找到第一个大于match_start的起始索引,减一就是目标单词的位置 i = bisect.bisect_right(start_indices, match_start) - 1 if i >= 0: start, end, word = word_positions[i] # 确保匹配索引确实在这个单词范围内(处理边界异常) if start <= match_start < end: if i > 0: print(f"前一个单词:{word_positions[i-1][2]}") if i + 2 < len(word_positions): print(f"后一个单词:{word_positions[i+2][2]}")
这种方法只需要一次预处理,后续的查找操作都是O(logn)级别的,比反复调用正则高效太多,非常适合大文件场景。
内容的提问来源于stack exchange,提问作者cryptojesus
相关产品推荐
相关产品推荐

