NLTK中如何正确统计各文本搭配的Conditional Frequency Distribution?
Hey there! Let's break down where your code went wrong and how to get the actual collocation occurrence counts across your texts.
What's Wrong with Your Current Code
The issue in your script is that you're not actually counting how many times each collocation appears in individual texts. Instead, you're creating a pairing between every text name and every collocation from the entire corpus—that's why every entry shows up as 1.
Let's look at the problematic line:
cfd = nltk.ConditionalFreqDist( (textname, collocation) for textname in eng_corpus.fileids() for collocation in Text(eng_corpus.words()).collocation_list(num=100))
Text(eng_corpus.words())uses all words from your entire corpus, not the current text in the loop.- You're just generating a tuple of (textname, collocation) for every possible pair, not tracking real occurrences.
The Correct Approach
We need to:
- Get the list of collocations we want to track (from the full corpus, as you did)
- For each text, generate its word pairs (bigrams)
- Count how many times each target collocation appears in those bigrams
- Feed those counts into the ConditionalFreqDist
Here's the fixed code:
import nltk from nltk.corpus import PlaintextCorpusReader # Assuming your corpus setup is already done (eng_corpus, tengc_low) eng_corpus_root = 'D:\\Corpus\\EN' eng_corpus = PlaintextCorpusReader(eng_corpus_root, '.*') # Step 1: Get target collocations and convert them to bigram tuples # (Collocation list returns strings like 'hong kong'; we need tuples ('hong', 'kong') to match bigrams) target_collocations = tengc_low.collocation_list(num=100) collocation_bigrams = [tuple(colloc.split()) for colloc in target_collocations] # Step 2: Initialize ConditionalFreqDist and populate with real counts cfd = nltk.ConditionalFreqDist() for text_id in eng_corpus.fileids(): # Get the current text's words (lowercase to match tengc_low) text_words = [word.lower() for word in eng_corpus.words(text_id)] # Generate all bigrams for the text text_bigrams = nltk.bigrams(text_words) # Count bigram frequencies for the text bigram_freq = nltk.FreqDist(text_bigrams) # Update the ConditionalFreqDist with counts for each target collocation for colloc in collocation_bigrams: # Use get() to return 0 if the collocation doesn't appear in the text cfd[text_id][colloc] = bigram_freq.get(colloc, 0)
How to Check the Results
Once you've built the corrected cfd, you can inspect the data like this:
- To see counts for a specific text:
# Replace 'sample_text.txt' with one of your file IDs print(cfd['sample_text.txt']) - To create a table showing collocation counts across texts (for the first 5 collocations, for example):
cfd.tabulate(conditions=eng_corpus.fileids()[:3], samples=collocation_bigrams[:5])
A Quick Optimization
If your corpus is large, using nltk.FreqDist for each text is more efficient than calling list(text_bigrams).count(colloc) every time, since it precomputes all bigram counts once per text.
内容的提问来源于stack exchange,提问作者Gavrk

