基于词频出现次数的DOM与数据库词频字典相似度计算问询
考虑词频的词频字典相似度计算方法
针对你需要结合词频计算DOM与数据库条目相似度的需求,以下是几种实用的实现方案,替代原有的仅统计共同词数量的方法:
方法1:加权共同词匹配得分
直接针对共同词,根据词频匹配度赋予权重,求和得到总得分。匹配度用「最小词频/最大词频」计算,范围0-1,词频越接近,贡献的得分越高。
def weighted_match_score(dom_counts, db_counts): common_words = set(dom_counts.keys()) & set(db_counts.keys()) total_score = 0.0 for word in common_words: dom_freq = dom_counts[word] db_freq = db_counts[word] # 词频完全匹配得1分,差异越大得分越低 match_weight = min(dom_freq, db_freq) / max(dom_freq, db_freq) total_score += match_weight return total_score # 测试示例数据 word_count_dom = {"Peter": 10, "is": 6, "eating": 2} word_count_db = {"eating": 6, "is": 6, "Peter": 1, "breakfast": 1} print(weighted_match_score(word_count_dom, word_count_db)) # 输出:1.4333333333333333
示例中,"is"因词频完全匹配贡献1分,"eating"贡献约0.33分,"Peter"仅贡献0.1分,完全符合你需要的权重逻辑。
方法2:余弦相似度
将两个词频字典视为高维向量,计算向量间的余弦夹角,值越接近1说明整体词频分布越相似。适合需要考虑所有词(包括独有词)分布差异的场景。
import math def cosine_similarity(dom_counts, db_counts): all_words = set(dom_counts.keys()) | set(db_counts.keys()) # 构建统一维度的词频向量 dom_vec = [dom_counts.get(word, 0) for word in all_words] db_vec = [db_counts.get(word, 0) for word in all_words] # 计算点积与向量模长 dot_product = sum(a * b for a, b in zip(dom_vec, db_vec)) dom_norm = math.sqrt(sum(x**2 for x in dom_vec)) db_norm = math.sqrt(sum(x**2 for x in db_vec)) if dom_norm == 0 or db_norm == 0: return 0.0 return dot_product / (dom_norm * db_norm) print(cosine_similarity(word_count_dom, word_count_db)) # 输出:0.5727644793833071
方法3:加权Jaccard相似度
基于传统Jaccard相似度改进,用共同词的最小词频之和作为交集,所有词的最大词频之和作为并集,比值即为相似度,兼顾共同词匹配和独有词的影响。
def weighted_jaccard(dom_counts, db_counts): common_words = set(dom_counts.keys()) & set(db_counts.keys()) union_words = set(dom_counts.keys()) | set(db_counts.keys()) intersection_sum = sum(min(dom_counts[word], db_counts[word]) for word in common_words) union_sum = sum(max(dom_counts.get(word, 0), db_counts.get(word, 0)) for word in union_words) return intersection_sum / union_sum if union_sum != 0 else 0.0 print(weighted_jaccard(word_count_dom, word_count_db)) # 输出:0.391304347826087
方法选择建议
- 若仅聚焦共同词的词频匹配度,优先选加权共同词匹配得分,逻辑简单且贴合你的权重需求;
- 若需衡量整体词频分布的相似性(含独有词影响),选余弦相似度;
- 若想平衡共同词匹配与整体覆盖度,选加权Jaccard相似度。
内容的提问来源于stack exchange,提问作者Kokaninchen
相关产品推荐
相关产品推荐

