Python中长复杂字符串自定义拆分后获取各句子偏移量的实现方法
解决方案
直接使用re.finditer匹配句子获取原生偏移,或者使用带捕获分组的re.split同步计算分隔符长度,两种方案都可以完全解决偏移不准的问题。
方案1:正则直接匹配句子(推荐)
不需要手动计算任何长度,所有偏移由正则引擎从原字符串直接返回,100%准确:
import re example_string = "First sentence. Second one with multiple spaces before.\nThird line starts with newline. Last sentence." # 正则和你原有的分句逻辑完全对齐 # re.DOTALL 支持句子内包含换行的场景,不需要可以去掉 pattern = re.compile(r'.*?\.(?=\s+[a-zA-Z]|\Z)', re.DOTALL) offset_phrase = [] list_phrase = [] for match in pattern.finditer(example_string): s_start, s_end = match.span() # 去掉句子前导空白,和原re.split返回的句子内容完全一致 stripped_start = s_start + len(match.group()) - len(match.group().lstrip()) list_phrase.append(example_string[stripped_start:s_end]) offset_phrase.append((stripped_start, s_end)) # 验证准确性,不会抛出异常 for sent, (start, end) in zip(list_phrase, offset_phrase): assert example_string[start:end] == sent
方案2:复用原有split规则
如果不想调整分句正则,可给分隔符加捕获分组,拆分后同步累加句子和分隔符的长度计算偏移:
import re example_string = "First sentence. Second one.\nThird one." # 给原分隔符正则加括号变成捕获分组,split会同时返回句子和分隔符 parts = re.split(r'(?<=\.)(\s+)(?=[a-zA-Z])', example_string) list_phrase = [] offset_phrase = [] current_offset = 0 # 每两个元素为一组:句子+分隔符 for idx in range(0, len(parts), 2): sent = parts[idx] start = current_offset end = current_offset + len(sent) list_phrase.append(sent) offset_phrase.append((start, end)) # 累加分隔符长度,准备下一个句子的起始偏移 if idx + 1 < len(parts): current_offset = end + len(parts[idx+1])
两种方案的时间复杂度均为O(n),长文本处理效率和你原有方案一致,完全支持多空格、换行、制表符等任意长度的分隔符场景。
内容的提问来源于stack exchange,提问作者Erwin
相关产品推荐
相关产品推荐

