You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于词频出现次数的DOM与数据库词频字典相似度计算问询

考虑词频的词频字典相似度计算方法

针对你需要结合词频计算DOM与数据库条目相似度的需求,以下是几种实用的实现方案,替代原有的仅统计共同词数量的方法:

方法1:加权共同词匹配得分

直接针对共同词,根据词频匹配度赋予权重,求和得到总得分。匹配度用「最小词频/最大词频」计算,范围0-1,词频越接近,贡献的得分越高。

def weighted_match_score(dom_counts, db_counts):
    common_words = set(dom_counts.keys()) & set(db_counts.keys())
    total_score = 0.0
    for word in common_words:
        dom_freq = dom_counts[word]
        db_freq = db_counts[word]
        # 词频完全匹配得1分,差异越大得分越低
        match_weight = min(dom_freq, db_freq) / max(dom_freq, db_freq)
        total_score += match_weight
    return total_score

# 测试示例数据
word_count_dom = {"Peter": 10, "is": 6, "eating": 2}
word_count_db = {"eating": 6, "is": 6, "Peter": 1, "breakfast": 1}
print(weighted_match_score(word_count_dom, word_count_db))
# 输出:1.4333333333333333

示例中,"is"因词频完全匹配贡献1分,"eating"贡献约0.33分,"Peter"仅贡献0.1分,完全符合你需要的权重逻辑。

方法2:余弦相似度

将两个词频字典视为高维向量,计算向量间的余弦夹角,值越接近1说明整体词频分布越相似。适合需要考虑所有词(包括独有词)分布差异的场景。

import math

def cosine_similarity(dom_counts, db_counts):
    all_words = set(dom_counts.keys()) | set(db_counts.keys())
    # 构建统一维度的词频向量
    dom_vec = [dom_counts.get(word, 0) for word in all_words]
    db_vec = [db_counts.get(word, 0) for word in all_words]
    
    # 计算点积与向量模长
    dot_product = sum(a * b for a, b in zip(dom_vec, db_vec))
    dom_norm = math.sqrt(sum(x**2 for x in dom_vec))
    db_norm = math.sqrt(sum(x**2 for x in db_vec))
    
    if dom_norm == 0 or db_norm == 0:
        return 0.0
    return dot_product / (dom_norm * db_norm)

print(cosine_similarity(word_count_dom, word_count_db))
# 输出:0.5727644793833071

方法3:加权Jaccard相似度

基于传统Jaccard相似度改进,用共同词的最小词频之和作为交集,所有词的最大词频之和作为并集,比值即为相似度,兼顾共同词匹配和独有词的影响。

def weighted_jaccard(dom_counts, db_counts):
    common_words = set(dom_counts.keys()) & set(db_counts.keys())
    union_words = set(dom_counts.keys()) | set(db_counts.keys())
    
    intersection_sum = sum(min(dom_counts[word], db_counts[word]) for word in common_words)
    union_sum = sum(max(dom_counts.get(word, 0), db_counts.get(word, 0)) for word in union_words)
    
    return intersection_sum / union_sum if union_sum != 0 else 0.0

print(weighted_jaccard(word_count_dom, word_count_db))
# 输出:0.391304347826087

方法选择建议

  • 若仅聚焦共同词的词频匹配度,优先选加权共同词匹配得分,逻辑简单且贴合你的权重需求;
  • 若需衡量整体词频分布的相似性(含独有词影响),选余弦相似度;
  • 若想平衡共同词匹配与整体覆盖度,选加权Jaccard相似度。

内容的提问来源于stack exchange,提问作者Kokaninchen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 01:01:51