You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中从字符串索引获取完整单词及目标前后单词(规避正则)

不用正则获取匹配目标前后的单词

嘿,我完全理解你不想用正则处理大文件的想法——正则虽然灵活,但在处理海量文本时确实容易成为性能瓶颈。既然你已经通过正则拿到了每个匹配目标的起始索引,我们可以换用纯字符串操作和索引定位的方式来获取前后的单词,这样效率会高很多,而且也能满足你后续的其他需求。

核心思路

先预处理文本,把每个单词的起始索引、结束索引、单词内容都记录下来,形成一个有序的列表。之后针对每个匹配的起始索引,只需要在这个列表里快速定位到目标单词的位置,就能直接拿到前后的单词了。

具体实现代码

首先保留你原来的匹配逻辑,获取所有目标的起始索引:

import re

text = 'hi this is john my name is john im bad boy'
target = 'is john'

# 你原有的匹配代码,获取每个匹配的起始索引
target_pattern = target.replace(' ', '[\s\n]*')
target_re = re.compile(r'\b%s' % target_pattern, flags=re.I | re.X)
indices = [m.start() for m in target_re.finditer(text)]

然后预处理文本,生成单词的位置映射:

# 生成每个单词的(起始索引, 结束索引, 单词内容)列表
word_positions = []
current_idx = 0
words = text.split()

for word in words:
    # 找到当前单词在文本中的起始位置
    start = text.find(word, current_idx)
    end = start + len(word)
    word_positions.append((start, end, word))
    # 更新当前索引到单词结束后的下一个位置(跳过空格)
    current_idx = end + 1

接下来处理每个匹配索引,获取前后单词:

# 遍历每个匹配的起始索引
for match_start in indices:
    # 找到匹配目标所在的单词位置(这里match_start是"is"的起始索引)
    for i, (start, end, word) in enumerate(word_positions):
        if start <= match_start < end:
            # 前一个单词:当前单词的前一位
            if i > 0:
                prev_word = word_positions[i-1][2]
                print(f"匹配'{target}'的前一个单词:{prev_word}")
            # 后一个单词:因为"is john"占2个单词,所以取当前位置+2的单词
            if i + 2 < len(word_positions):
                next_word = word_positions[i+2][2]
                print(f"匹配'{target}'的后一个单词:{next_word}")
            break

性能优化:用二分查找快速定位

如果是处理超大文件,上面的遍历查找可以换成二分查找,进一步提升速度:

import bisect

# 提取所有单词的起始索引,用于二分查找
start_indices = [pos[0] for pos in word_positions]

for match_start in indices:
    # 用bisect找到第一个大于match_start的起始索引,减一就是目标单词的位置
    i = bisect.bisect_right(start_indices, match_start) - 1
    if i >= 0:
        start, end, word = word_positions[i]
        # 确保匹配索引确实在这个单词范围内(处理边界异常)
        if start <= match_start < end:
            if i > 0:
                print(f"前一个单词:{word_positions[i-1][2]}")
            if i + 2 < len(word_positions):
                print(f"后一个单词:{word_positions[i+2][2]}")

这种方法只需要一次预处理,后续的查找操作都是O(logn)级别的,比反复调用正则高效太多,非常适合大文件场景。

内容的提问来源于stack exchange,提问作者cryptojesus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:58:09