You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何实现包含子串匹配的句子交集计算?

实现考虑子串/缩写匹配的Jaccard相似度交集

默认的set.intersection仅支持精确字符串匹配,要将'Oxford'与'Oxf.'这类存在子串/缩写关联的词纳入交集,需要自定义模糊匹配逻辑,以下是具体实现方案:

1. 定义预处理与匹配函数

先对词做标准化处理(去除标点、统一大小写),再定义匹配规则(兼顾前缀/子串匹配,加入长度限制减少误匹配):

def preprocess_word(word):
    # 去除词末尾的标点,转小写统一格式
    return word.rstrip('.,!?').lower()

def is_matching(word1, word2):
    w1 = preprocess_word(word1)
    w2 = preprocess_word(word2)
    len1, len2 = len(w1), len(w2)
    
    # 空词直接跳过匹配
    if len1 == 0 or len2 == 0:
        return False
    
    # 限制长度差异,避免短词误匹配长词(比如"cat"和"category")
    min_len = min(len1, len2)
    max_len = max(len1, len2)
    if min_len / max_len < 0.5:
        return False
    
    # 判断是否为前缀匹配或子串匹配
    if len1 <= len2:
        return w2.startswith(w1) or w1 in w2
    else:
        return w1.startswith(w2) or w2 in w1

2. 找出模糊匹配的交集

遍历两个词列表,收集所有满足匹配规则的词对,再整理去重得到交集:

item1 = 'She went to a restaurant on Oxford Street'.split(' ')
item2 = 'She went to an Italian restaurant on Oxf. Street'.split(' ')

# 收集所有匹配的词对
matched_pairs = []
for w1 in item1:
    for w2 in item2:
        if is_matching(w1, w2):
            matched_pairs.append((w1, w2))

# 生成去重后的交集集合(保留原词形式)
intersection = set()
for w1, w2 in matched_pairs:
    intersection.add(w1)
    intersection.add(w2)

print(intersection)
# 输出:{'She', 'Street', 'on', 'restaurant', 'to', 'went', 'Oxford', 'Oxf.'}

如果希望交集只保留每个匹配组的标准词(比如较长的原词),可以调整为:

intersection = set()
seen_groups = set()

for w1, w2 in matched_pairs:
    # 用预处理后的词生成唯一组标识,避免重复处理
    group_key = frozenset({preprocess_word(w1), preprocess_word(w2)})
    if group_key not in seen_groups:
        seen_groups.add(group_key)
        # 选择较长的词作为组代表
        intersection.add(w1 if len(w1) > len(w2) else w2)

print(intersection)
# 输出:{'She', 'Street', 'on', 'restaurant', 'to', 'went', 'Oxford'}

3. 计算考虑模糊匹配的Jaccard相似度

Jaccard相似度公式为交集大小 / 并集大小,这里需要基于匹配等价组计算:

from collections import defaultdict

def get_equivalence_groups(word_list, match_func):
    """把匹配的词归为同一等价组"""
    groups = []
    used_indices = set()
    
    for idx, word in enumerate(word_list):
        if idx in used_indices:
            continue
        current_group = {word}
        for other_idx, other_word in enumerate(word_list):
            if other_idx not in used_indices and match_func(word, other_word):
                current_group.add(other_word)
                used_indices.add(other_idx)
        groups.append(current_group)
    return groups

# 合并两个列表的词,生成所有等价组
all_words = item1 + item2
all_groups = get_equivalence_groups(all_words, is_matching)

# 计算交集组数(同时出现在两个原列表中的组)
intersection_count = 0
for group in all_groups:
    has_from_item1 = any(w in item1 for w in group)
    has_from_item2 = any(w in item2 for w in group)
    if has_from_item1 and has_from_item2:
        intersection_count += 1

# 并集组数即总等价组数
union_count = len(all_groups)

# 计算Jaccard相似度
jaccard_similarity = intersection_count / union_count
print(jaccard_similarity)

注意事项

  • 预处理规则可按需调整:比如保留大小写(去掉.lower())、处理更多特殊字符。
  • 匹配规则可优化:如果有特定缩写词典,可先替换缩写再计算,比通用子串匹配更精准。
  • 长度限制阈值可灵活调整:若允许更短的缩写,可将0.5调低。

内容的提问来源于stack exchange,提问作者user19495470

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 12:20:28