You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中高效检索大文本中的相似子串?

解决Python中大文本相似子串的检索问题

核心思路

由于目标子串是原语料通过re.sub("[^a-zA-Z]", " ", corpus)预处理得到的,二者的核心单词序列完全一致,只是原语料多了标点、缩写符号等非字母内容。因此可以通过对齐原语料的单词与原位置映射,匹配子串的单词序列来精准定位原语料中的起止索引。

具体实现代码

import re

corpus = """very quick service, polite workers(cory, i think that's his name), i basically just drove there and got a quote(which seems to be very fair priced), then dropped off my car 4 days later(because they were fully booked until then), then i dropped off my car on my appointment day, then the same day the shop called me and notified me that the the job is done i can go pickup my car. when i go checked out my car i was amazed by the job they've done to it, and they even gave that dirty car a wash( prob even waxed it or coated it, cuz it was shiny as hell), tires shine, mats were vacuumed too. i gave them a dirty, broken car, they gave me back a what seems like a brand new car. i'm happy with the result, and i will def have all my car's work done by this place from now."""
substring = """until then then i dropped off my car on my appointment day then the same day the shop called me and notified me that the the job is done i can go pickup my car when i go checked out my car i was amazed by the job they ve done to it and they even gave that dirty car a wash prob even waxed it or coated it cuz it was shiny as hell tires shine mats were vacuumed too i gave them a dirty broken car they gave me back a what seems like a brand new car i m happy with the result and i will def have all my car s work done by this place from now"""

def find_substring_indices(corpus, substring):
    # 提取原语料中所有字母单词,并记录每个单词在原语料的起止索引
    corpus_word_info = []
    current_word = []
    start_idx = None
    for idx, char in enumerate(corpus):
        if char.isalpha():
            if not current_word:
                start_idx = idx
            current_word.append(char)
        else:
            if current_word:
                corpus_word_info.append(
                    (''.join(current_word), start_idx, idx)
                )
                current_word = []
    # 处理语料末尾的单词
    if current_word:
        corpus_word_info.append(
            (''.join(current_word), start_idx, len(corpus))
        )
    
    # 将目标子串拆分为单词列表
    target_words = substring.split()
    target_len = len(target_words)

    # 滑动窗口匹配单词序列
    for i in range(len(corpus_word_info) - target_len + 1):
        match_flag = True
        for j in range(target_len):
            # 忽略大小写匹配,可根据需求调整为严格匹配
            if corpus_word_info[i+j][0].lower() != target_words[j].lower():
                match_flag = False
                break
        if match_flag:
            # 返回原语料中对应片段的起止索引
            return (corpus_word_info[i][1], corpus_word_info[i+target_len-1][2])
    return None

# 测试执行
match_indices = find_substring_indices(corpus, substring)
print(f"匹配到的起止索引: {match_indices}")
if match_indices:
    print("验证匹配片段:")
    print(corpus[match_indices[0]:match_indices[1]])

代码说明

  1. 原语料单词映射:遍历原语料,提取所有由字母组成的单词,同时记录每个单词在原语料中的起始和结束索引,建立单词与原位置的对应关系。
  2. 子串拆分:将目标子串按空格拆分为单词列表,由于子串是原语料预处理后的结果,二者的单词序列完全一致。
  3. 滑动窗口匹配:在原语料的单词列表中滑动匹配目标单词序列,找到匹配后返回对应单词在原语料中的起止索引,即为目标片段的位置。

内容的提问来源于stack exchange,提问作者user_12

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 05:35:22