You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python3中比较NLTK FreqDist生成的频率分布并输出数值结果?

对比语料库与单个文本的FreqDist分布

没问题,我来帮你一步步实现这个需求!核心思路是先为每个单个文本生成独立的FreqDist,然后通过统计指标、高频词对比、分布相似度计算等方式来和语料库的全局分布做比较。

第一步:生成单个文本的频率分布

首先,你需要在遍历语料库的同时,为每个单独的文本创建FreqDist并保存起来,这样后续才能逐一对比。修改你的现有代码如下:

from nltk.probability import FreqDist

# 初始化存储变量
corpus_tokens = []
document_fdists = []  # 存储每个单个文本的FreqDist对象

# 遍历所有文档,同时生成全局和单个文档的频率分布
for document in documents_set:
    corpus_tokens.extend(document)
    # 为当前文档生成独立的FreqDist并添加到列表
    doc_fdist = FreqDist(document)
    document_fdists.append(doc_fdist)

# 生成语料库全局的频率分布
corpus_fdist = FreqDist(corpus_tokens)

第二步:实现对比逻辑与数值输出

接下来我们写一个对比函数,包含几种实用的对比维度,你可以根据需求调整或扩展:

import math

def compare_single_to_corpus(corpus_fdist, doc_fdist, doc_index):
    print(f"=== 第 {doc_index+1} 个文本 vs 语料库频率分布对比 ===")
    
    # 1. 输出两者的Top N高频词
    top_n = 10
    print(f"\n📊 语料库Top {top_n}高频词:")
    for word, count in corpus_fdist.most_common(top_n):
        print(f"  {word}: {count}")
    
    print(f"\n📄 当前文本Top {top_n}高频词:")
    for word, count in doc_fdist.most_common(top_n):
        print(f"  {word}: {count}")
    
    # 2. 统计高频词重叠情况(取Top 50)
    corpus_top_50 = set(word for word, _ in corpus_fdist.most_common(50))
    doc_top_50 = set(word for word, _ in doc_fdist.most_common(50))
    overlap = corpus_top_50 & doc_top_50
    print(f"\n🔄 Top 50高频词重叠数量: {len(overlap)}")
    print(f"   重叠词列表: {sorted(overlap)}")
    
    # 3. 当前文本独有的高频词(不在语料库Top 50)
    unique_top = [word for word, _ in doc_fdist.most_common(20) if word not in corpus_top_50][:5]
    print(f"\n🆕 当前文本独有的Top 5高频词: {unique_top}")
    
    # 4. 核心词的频率占比对比
    if doc_fdist.most_common(1):
        top_word = doc_fdist.most_common(1)[0][0]
        corpus_total = sum(corpus_fdist.values())
        doc_total = sum(doc_fdist.values())
        
        corpus_freq = corpus_fdist[top_word] / corpus_total if corpus_total > 0 else 0
        doc_freq = doc_fdist[top_word] / doc_total if doc_total > 0 else 0
        
        print(f"\n🔍 当前文本核心词 '{top_word}' 频率对比:")
        print(f"   语料库中占比: {corpus_freq:.4f}")
        print(f"   当前文本中占比: {doc_freq:.4f}")
        if corpus_freq > 0:
            print(f"   相对语料库的频率倍数: {doc_freq / corpus_freq:.2f}")
        else:
            print(f"   该词未出现在语料库中")
    
    # 5. KL散度(衡量两个分布的差异,值越小越相似)
    def calculate_kl(p, q):
        kl_sum = 0.0
        all_words = set(p.keys()).union(set(q.keys()))
        p_total = sum(p.values())
        q_total = sum(q.values())
        for word in all_words:
            p_prob = p[word] / p_total if p_total > 0 else 0
            q_prob = q[word] / q_total if q_total > 0 else 0
            if p_prob > 0 and q_prob > 0:
                kl_sum += p_prob * math.log(p_prob / q_prob)
        return kl_sum
    
    kl_val = calculate_kl(corpus_fdist, doc_fdist)
    print(f"\n📏 KL散度(分布差异值): {kl_val:.4f}")

第三步:遍历所有文本执行对比

最后,遍历之前保存的单个文本FreqDist列表,逐个调用对比函数:

# 遍历所有文档进行对比
for idx, doc_fdist in enumerate(document_fdists):
    compare_single_to_corpus(corpus_fdist, doc_fdist, idx)
    print("\n" + "-"*60 + "\n")

补充说明

  • 如果你需要更直观的相似度衡量,可以实现余弦相似度:将两个频率分布转化为向量,计算点积除以模长的乘积,结果越接近1说明分布越相似。
  • 可以根据需求调整对比的维度,比如添加低频词差异、特定词性的频率对比等。

内容的提问来源于stack exchange,提问作者ug chauhan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 04:02:02