Python中如何实现包含子串匹配的句子交集计算?
实现考虑子串/缩写匹配的Jaccard相似度交集
默认的set.intersection仅支持精确字符串匹配,要将'Oxford'与'Oxf.'这类存在子串/缩写关联的词纳入交集,需要自定义模糊匹配逻辑,以下是具体实现方案:
1. 定义预处理与匹配函数
先对词做标准化处理(去除标点、统一大小写),再定义匹配规则(兼顾前缀/子串匹配,加入长度限制减少误匹配):
def preprocess_word(word): # 去除词末尾的标点,转小写统一格式 return word.rstrip('.,!?').lower() def is_matching(word1, word2): w1 = preprocess_word(word1) w2 = preprocess_word(word2) len1, len2 = len(w1), len(w2) # 空词直接跳过匹配 if len1 == 0 or len2 == 0: return False # 限制长度差异,避免短词误匹配长词(比如"cat"和"category") min_len = min(len1, len2) max_len = max(len1, len2) if min_len / max_len < 0.5: return False # 判断是否为前缀匹配或子串匹配 if len1 <= len2: return w2.startswith(w1) or w1 in w2 else: return w1.startswith(w2) or w2 in w1
2. 找出模糊匹配的交集
遍历两个词列表,收集所有满足匹配规则的词对,再整理去重得到交集:
item1 = 'She went to a restaurant on Oxford Street'.split(' ') item2 = 'She went to an Italian restaurant on Oxf. Street'.split(' ') # 收集所有匹配的词对 matched_pairs = [] for w1 in item1: for w2 in item2: if is_matching(w1, w2): matched_pairs.append((w1, w2)) # 生成去重后的交集集合(保留原词形式) intersection = set() for w1, w2 in matched_pairs: intersection.add(w1) intersection.add(w2) print(intersection) # 输出:{'She', 'Street', 'on', 'restaurant', 'to', 'went', 'Oxford', 'Oxf.'}
如果希望交集只保留每个匹配组的标准词(比如较长的原词),可以调整为:
intersection = set() seen_groups = set() for w1, w2 in matched_pairs: # 用预处理后的词生成唯一组标识,避免重复处理 group_key = frozenset({preprocess_word(w1), preprocess_word(w2)}) if group_key not in seen_groups: seen_groups.add(group_key) # 选择较长的词作为组代表 intersection.add(w1 if len(w1) > len(w2) else w2) print(intersection) # 输出:{'She', 'Street', 'on', 'restaurant', 'to', 'went', 'Oxford'}
3. 计算考虑模糊匹配的Jaccard相似度
Jaccard相似度公式为交集大小 / 并集大小,这里需要基于匹配等价组计算:
from collections import defaultdict def get_equivalence_groups(word_list, match_func): """把匹配的词归为同一等价组""" groups = [] used_indices = set() for idx, word in enumerate(word_list): if idx in used_indices: continue current_group = {word} for other_idx, other_word in enumerate(word_list): if other_idx not in used_indices and match_func(word, other_word): current_group.add(other_word) used_indices.add(other_idx) groups.append(current_group) return groups # 合并两个列表的词,生成所有等价组 all_words = item1 + item2 all_groups = get_equivalence_groups(all_words, is_matching) # 计算交集组数(同时出现在两个原列表中的组) intersection_count = 0 for group in all_groups: has_from_item1 = any(w in item1 for w in group) has_from_item2 = any(w in item2 for w in group) if has_from_item1 and has_from_item2: intersection_count += 1 # 并集组数即总等价组数 union_count = len(all_groups) # 计算Jaccard相似度 jaccard_similarity = intersection_count / union_count print(jaccard_similarity)
注意事项
- 预处理规则可按需调整:比如保留大小写(去掉
.lower())、处理更多特殊字符。 - 匹配规则可优化:如果有特定缩写词典,可先替换缩写再计算,比通用子串匹配更精准。
- 长度限制阈值可灵活调整:若允许更短的缩写,可将
0.5调低。
内容的提问来源于stack exchange,提问作者user19495470
相关产品推荐
相关产品推荐

