如何按词频降序排序目标词并统计同句关联词(无额外模块)
Got it, let's work through this step by step using only Python's built-in tools. I'll cover fixing the sorting issue and implementing the co-occurrence word stats you need.
Step 1: Load Common Words & Process Sample Text
First, we'll load the common words into a set for fast lookup (way quicker than a list for membership checks), then split the sample text into sentences and strip out common words while preserving sentence structure.
# Load common words into a case-insensitive set with open('common.txt', 'r', encoding='utf-8') as f: common_words = set(word.strip().lower() for word in f if word.strip()) # Load sample text and split into sentences with open('sample.txt', 'r', encoding='utf-8') as f: sample_text = f.read() # Split sentences by periods, clean up empty/whitespace-only entries sentences = [sent.strip() for sent in sample_text.split('.') if sent.strip()] # Optional: Helper to clean punctuation from words (avoids counting "apple" vs "apple," as separate) def clean_word(word): punctuation = '.,!?;:()[]{}"\'' return ''.join(char for char in word if char not in punctuation).lower() # Process each sentence: remove common words, keep valid words processed_sentences = [] for sent in sentences: # Split into cleaned words, filter out common words and empty strings words = [ clean_word(word) for word in sent.split() if clean_word(word) and clean_word(word) not in common_words ] if words: # Skip sentences that become empty after filtering processed_sentences.append(words)
Step 2: Count Target Words & Co-Occurring Words
We'll use two dictionaries: one for tracking target word frequencies, and another to log how often each word co-occurs with the target in the same sentence.
# Initialize counters target_word_counts = {} cooccurrence_counts = {} for word_list in processed_sentences: # Iterate over each word as the target for idx, target_word in enumerate(word_list): # Update target word's total count target_word_counts[target_word] = target_word_counts.get(target_word, 0) + 1 # Get all other words in the sentence as co-occurrences co_words = word_list[:idx] + word_list[idx+1:] # Update co-occurrence counts for this target if target_word not in cooccurrence_counts: cooccurrence_counts[target_word] = {} for co_word in co_words: cooccurrence_counts[target_word][co_word] = cooccurrence_counts[target_word].get(co_word, 0) + 1
Step 3: Sort Target Words by Frequency (Descending)
Use Python's built-in sorted() function to sort the target words by their count in descending order. We'll sort the dictionary items using the count as the key.
# Sort target words by frequency (highest first) sorted_targets = sorted(target_word_counts.items(), key=lambda x: x[1], reverse=True)
Step 4: Print or Export Results
Here's how you can display the sorted targets along with their co-occurring words (also sorted by frequency for clarity):
# Display sorted results print("Sorted Target Words (by frequency):") for word, count in sorted_targets: print(f"- {word}: {count} occurrences") # Show top co-occurring words (sorted by frequency) if word in cooccurrence_counts: sorted_coocurrences = sorted(cooccurrence_counts[word].items(), key=lambda x: x[1], reverse=True) print(" Top co-occurring words:") for co_word, co_count in sorted_coocurrences[:5]: # Show top 5, adjust as needed print(f" • {co_word}: {co_count} times") print("---")
Quick Notes on Edge Cases
- Case Insensitivity: All words are converted to lowercase to avoid counting "Apple" and "apple" as separate entries.
- Punctuation Handling: The optional
clean_wordfunction strips punctuation to ensure consistent counting across variants like "apple" and "apple,". - Empty Sentences: Sentences that become empty after removing common words are skipped to avoid skewing stats.
内容的提问来源于stack exchange,提问作者Jim Ye

