如何在Python中高效检索大文本中的相似子串?
解决Python中大文本相似子串的检索问题
核心思路
由于目标子串是原语料通过re.sub("[^a-zA-Z]", " ", corpus)预处理得到的,二者的核心单词序列完全一致,只是原语料多了标点、缩写符号等非字母内容。因此可以通过对齐原语料的单词与原位置映射,匹配子串的单词序列来精准定位原语料中的起止索引。
具体实现代码
import re corpus = """very quick service, polite workers(cory, i think that's his name), i basically just drove there and got a quote(which seems to be very fair priced), then dropped off my car 4 days later(because they were fully booked until then), then i dropped off my car on my appointment day, then the same day the shop called me and notified me that the the job is done i can go pickup my car. when i go checked out my car i was amazed by the job they've done to it, and they even gave that dirty car a wash( prob even waxed it or coated it, cuz it was shiny as hell), tires shine, mats were vacuumed too. i gave them a dirty, broken car, they gave me back a what seems like a brand new car. i'm happy with the result, and i will def have all my car's work done by this place from now.""" substring = """until then then i dropped off my car on my appointment day then the same day the shop called me and notified me that the the job is done i can go pickup my car when i go checked out my car i was amazed by the job they ve done to it and they even gave that dirty car a wash prob even waxed it or coated it cuz it was shiny as hell tires shine mats were vacuumed too i gave them a dirty broken car they gave me back a what seems like a brand new car i m happy with the result and i will def have all my car s work done by this place from now""" def find_substring_indices(corpus, substring): # 提取原语料中所有字母单词,并记录每个单词在原语料的起止索引 corpus_word_info = [] current_word = [] start_idx = None for idx, char in enumerate(corpus): if char.isalpha(): if not current_word: start_idx = idx current_word.append(char) else: if current_word: corpus_word_info.append( (''.join(current_word), start_idx, idx) ) current_word = [] # 处理语料末尾的单词 if current_word: corpus_word_info.append( (''.join(current_word), start_idx, len(corpus)) ) # 将目标子串拆分为单词列表 target_words = substring.split() target_len = len(target_words) # 滑动窗口匹配单词序列 for i in range(len(corpus_word_info) - target_len + 1): match_flag = True for j in range(target_len): # 忽略大小写匹配,可根据需求调整为严格匹配 if corpus_word_info[i+j][0].lower() != target_words[j].lower(): match_flag = False break if match_flag: # 返回原语料中对应片段的起止索引 return (corpus_word_info[i][1], corpus_word_info[i+target_len-1][2]) return None # 测试执行 match_indices = find_substring_indices(corpus, substring) print(f"匹配到的起止索引: {match_indices}") if match_indices: print("验证匹配片段:") print(corpus[match_indices[0]:match_indices[1]])
代码说明
- 原语料单词映射:遍历原语料,提取所有由字母组成的单词,同时记录每个单词在原语料中的起始和结束索引,建立单词与原位置的对应关系。
- 子串拆分:将目标子串按空格拆分为单词列表,由于子串是原语料预处理后的结果,二者的单词序列完全一致。
- 滑动窗口匹配:在原语料的单词列表中滑动匹配目标单词序列,找到匹配后返回对应单词在原语料中的起止索引,即为目标片段的位置。
内容的提问来源于stack exchange,提问作者user_12
相关产品推荐
相关产品推荐

