Python如何比对多个字符串并提取所有字符串共有的子串
多字符串公共连续单词词组查找实现
核心思路
目标是提取所有输入字符串中都存在的最长连续单词组合,逻辑如下:
- 先把所有输入字符串拆分为独立的单词列表
- 选择单词数最少的列表作为匹配基准(公共词组的长度不可能超过最短字符串的单词总数,能大幅减少无效计算)
- 从最长的可能词组长度开始向下遍历,逐个生成候选词组
- 只要找到一个候选词组在所有字符串的单词列表中都以连续形式存在,直接返回该词组,就是符合要求的最长公共结果
纯Python实现(无第三方依赖)
普通空格分隔的简单文本直接用以下代码即可,不需要额外安装工具:
def find_longest_common_phrase(input_strings): # 拆分所有字符串为单词列表,过滤空字符串 word_lists = [s.split() for s in input_strings if s.strip()] if not word_lists: return "" # 取最短单词列表作为匹配基准 base_words = min(word_lists, key=len) max_possible_len = len(base_words) # 从最长词组长度开始往下匹配 for phrase_length in range(max_possible_len, 0, -1): # 生成基准列表中所有当前长度的连续词组 for start_idx in range(max_possible_len - phrase_length + 1): candidate = tuple(base_words[start_idx : start_idx + phrase_length]) # 检查候选词组是否在所有单词列表中存在 is_common = True for wl in word_lists: found = False wl_length = len(wl) # 遍历当前词列表找匹配的连续序列 for i in range(wl_length - phrase_length + 1): if tuple(wl[i : i + phrase_length]) == candidate: found = True break if not found: is_common = False break if is_common: return " ".join(candidate) return "" # 测试示例 few_strings = ('this is foo bar', 'this is not a foo bar', 'some other foo bar here') result = find_longest_common_phrase(few_strings) print(result)
运行上述代码输出结果为foo bar,和示例预期一致。
复杂文本适配(nltk优化分词)
如果待处理文本包含标点、大小写不统一,直接用split()会出现匹配误差,可以用nltk的分词工具做预处理:
import nltk nltk.download('punkt', quiet=True) from nltk.tokenize import word_tokenize def text_preprocess(text): # 统一转小写,分词后过滤纯标点符号 words = word_tokenize(text.lower()) return [word for word in words if word.isalnum()]
只需要把原函数中s.split()替换为text_preprocess(s),就可以适配带标点、大小写混杂的常规英文文本场景。
内容的提问来源于stack exchange,提问作者stkvtflw
相关产品推荐
相关产品推荐

