You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NLTK中如何正确统计各文本搭配的Conditional Frequency Distribution?

Fixing ConditionalFreqDist for Collocation Counts Per Text

Hey there! Let's break down where your code went wrong and how to get the actual collocation occurrence counts across your texts.

What's Wrong with Your Current Code

The issue in your script is that you're not actually counting how many times each collocation appears in individual texts. Instead, you're creating a pairing between every text name and every collocation from the entire corpus—that's why every entry shows up as 1.

Let's look at the problematic line:

cfd = nltk.ConditionalFreqDist( (textname, collocation) for textname in eng_corpus.fileids() for collocation in Text(eng_corpus.words()).collocation_list(num=100))
  • Text(eng_corpus.words()) uses all words from your entire corpus, not the current text in the loop.
  • You're just generating a tuple of (textname, collocation) for every possible pair, not tracking real occurrences.

The Correct Approach

We need to:

  1. Get the list of collocations we want to track (from the full corpus, as you did)
  2. For each text, generate its word pairs (bigrams)
  3. Count how many times each target collocation appears in those bigrams
  4. Feed those counts into the ConditionalFreqDist

Here's the fixed code:

import nltk
from nltk.corpus import PlaintextCorpusReader

# Assuming your corpus setup is already done (eng_corpus, tengc_low)
eng_corpus_root = 'D:\\Corpus\\EN'
eng_corpus = PlaintextCorpusReader(eng_corpus_root, '.*')

# Step 1: Get target collocations and convert them to bigram tuples
# (Collocation list returns strings like 'hong kong'; we need tuples ('hong', 'kong') to match bigrams)
target_collocations = tengc_low.collocation_list(num=100)
collocation_bigrams = [tuple(colloc.split()) for colloc in target_collocations]

# Step 2: Initialize ConditionalFreqDist and populate with real counts
cfd = nltk.ConditionalFreqDist()

for text_id in eng_corpus.fileids():
    # Get the current text's words (lowercase to match tengc_low)
    text_words = [word.lower() for word in eng_corpus.words(text_id)]
    # Generate all bigrams for the text
    text_bigrams = nltk.bigrams(text_words)
    # Count bigram frequencies for the text
    bigram_freq = nltk.FreqDist(text_bigrams)
    
    # Update the ConditionalFreqDist with counts for each target collocation
    for colloc in collocation_bigrams:
        # Use get() to return 0 if the collocation doesn't appear in the text
        cfd[text_id][colloc] = bigram_freq.get(colloc, 0)

How to Check the Results

Once you've built the corrected cfd, you can inspect the data like this:

  • To see counts for a specific text:
    # Replace 'sample_text.txt' with one of your file IDs
    print(cfd['sample_text.txt'])
    
  • To create a table showing collocation counts across texts (for the first 5 collocations, for example):
    cfd.tabulate(conditions=eng_corpus.fileids()[:3], samples=collocation_bigrams[:5])
    

A Quick Optimization

If your corpus is large, using nltk.FreqDist for each text is more efficient than calling list(text_bigrams).count(colloc) every time, since it precomputes all bigram counts once per text.

内容的提问来源于stack exchange,提问作者Gavrk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:55:57