You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求编写Python程序:统计文件前5高频首词,排除含DM/RT的行

How to Count Top 5 Most Frequent First Words (Excluding Lines with "DM"/"RT" After First Word)

Hey there! Let's walk through this step by step since you're new to coding—we'll build the solution from your initial code snippet, no jargon overload.

Step 1: Use a Safer Way to Open Files

First, instead of just open(), we'll use a with statement. It automatically closes the file when we're done, which avoids accidental resource leaks. Your initial file path is fine, we'll plug that right in.

Step 2: Filter Lines & Collect Valid First Words

We need to:

  • Skip empty lines (they'll cause errors if we try to grab the first word)
  • Skip lines where the second word is "DM" or "RT"
  • Collect the first word from all remaining valid lines

Step 3: Count Frequencies & Grab Top 5

Python has a built-in tool called Counter (from the collections module) that makes counting word frequencies super easy. We'll use it to tally up first words, then pull the top 5 most common ones.


Full Code with Clear Comments

# Import the Counter tool to simplify frequency counting
from collections import Counter

# Use 'with' to safely open and read your text file
with open("C:/Users/Joe Simpleton/Desktop/talking.txt", "r") as f:
    # Initialize an empty Counter to track how often each first word appears
    first_word_counts = Counter()
    
    # Loop through every line in the file
    for line in f:
        # Remove extra whitespace (like newlines or leading/trailing spaces) from the line
        cleaned_line = line.strip()
        
        # Skip empty lines to avoid errors later
        if not cleaned_line:
            continue
        
        # Split the line into a list of individual words (splits on spaces by default)
        words = cleaned_line.split()
        
        # Check if the line has at least 2 words, and if the second is "DM" or "RT"
        if len(words) >= 2 and words[1] in ("DM", "RT"):
            # Skip this line if it matches our exclusion rule
            continue
        
        # Grab the first word from the valid line
        first_word = words[0]
        # Optional: Uncomment below to make counting case-insensitive (e.g., "Hello" and "hello" count as the same)
        # first_word = first_word.lower()
        
        # Add this first word to our counter (increment its count by 1)
        first_word_counts[first_word] += 1

# Get the top 5 most frequent first words
top_5_words = first_word_counts.most_common(5)

# Print the result in a readable format
print("Top 5 most frequent first words:")
for word, count in top_5_words:
    print(f"- {word}: {count} times")

Quick Explanation of Key Parts

  • from collections import Counter: Pulls in a pre-built tool that handles counting repetitive items perfectly.
  • with open(...) as f: Opens your file and ensures it gets closed automatically, even if something goes wrong.
  • line.strip(): Cleans up messy line endings so we don't accidentally count empty strings as words.
  • The if len(words) >=2 ... check: Prevents crashes from single-word lines, and skips lines that meet your exclusion criteria.
  • most_common(5): Gives us a list of the 5 words that appear most often, paired with their total counts.

Optional Tweak

If you want to treat "Hello" and "hello" as the same word, uncomment the line first_word = first_word.lower()—this converts all first words to lowercase before counting.

内容的提问来源于stack exchange,提问作者johnson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:39:19