You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按词频降序排序目标词并统计同句关联词(无额外模块)

Pure Python Solution (No Extra Modules)

Got it, let's work through this step by step using only Python's built-in tools. I'll cover fixing the sorting issue and implementing the co-occurrence word stats you need.

Step 1: Load Common Words & Process Sample Text

First, we'll load the common words into a set for fast lookup (way quicker than a list for membership checks), then split the sample text into sentences and strip out common words while preserving sentence structure.

# Load common words into a case-insensitive set
with open('common.txt', 'r', encoding='utf-8') as f:
    common_words = set(word.strip().lower() for word in f if word.strip())

# Load sample text and split into sentences
with open('sample.txt', 'r', encoding='utf-8') as f:
    sample_text = f.read()

# Split sentences by periods, clean up empty/whitespace-only entries
sentences = [sent.strip() for sent in sample_text.split('.') if sent.strip()]

# Optional: Helper to clean punctuation from words (avoids counting "apple" vs "apple," as separate)
def clean_word(word):
    punctuation = '.,!?;:()[]{}"\''
    return ''.join(char for char in word if char not in punctuation).lower()

# Process each sentence: remove common words, keep valid words
processed_sentences = []
for sent in sentences:
    # Split into cleaned words, filter out common words and empty strings
    words = [
        clean_word(word) 
        for word in sent.split() 
        if clean_word(word) and clean_word(word) not in common_words
    ]
    if words:  # Skip sentences that become empty after filtering
        processed_sentences.append(words)

Step 2: Count Target Words & Co-Occurring Words

We'll use two dictionaries: one for tracking target word frequencies, and another to log how often each word co-occurs with the target in the same sentence.

# Initialize counters
target_word_counts = {}
cooccurrence_counts = {}

for word_list in processed_sentences:
    # Iterate over each word as the target
    for idx, target_word in enumerate(word_list):
        # Update target word's total count
        target_word_counts[target_word] = target_word_counts.get(target_word, 0) + 1
        
        # Get all other words in the sentence as co-occurrences
        co_words = word_list[:idx] + word_list[idx+1:]
        
        # Update co-occurrence counts for this target
        if target_word not in cooccurrence_counts:
            cooccurrence_counts[target_word] = {}
        
        for co_word in co_words:
            cooccurrence_counts[target_word][co_word] = cooccurrence_counts[target_word].get(co_word, 0) + 1

Step 3: Sort Target Words by Frequency (Descending)

Use Python's built-in sorted() function to sort the target words by their count in descending order. We'll sort the dictionary items using the count as the key.

# Sort target words by frequency (highest first)
sorted_targets = sorted(target_word_counts.items(), key=lambda x: x[1], reverse=True)

Step 4: Print or Export Results

Here's how you can display the sorted targets along with their co-occurring words (also sorted by frequency for clarity):

# Display sorted results
print("Sorted Target Words (by frequency):")
for word, count in sorted_targets:
    print(f"- {word}: {count} occurrences")
    
    # Show top co-occurring words (sorted by frequency)
    if word in cooccurrence_counts:
        sorted_coocurrences = sorted(cooccurrence_counts[word].items(), key=lambda x: x[1], reverse=True)
        print("  Top co-occurring words:")
        for co_word, co_count in sorted_coocurrences[:5]:  # Show top 5, adjust as needed
            print(f"    • {co_word}: {co_count} times")
    print("---")

Quick Notes on Edge Cases

  • Case Insensitivity: All words are converted to lowercase to avoid counting "Apple" and "apple" as separate entries.
  • Punctuation Handling: The optional clean_word function strips punctuation to ensure consistent counting across variants like "apple" and "apple,".
  • Empty Sentences: Sentences that become empty after removing common words are skipped to avoid skewing stats.

内容的提问来源于stack exchange,提问作者Jim Ye

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:23:28