You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用os.walk与NLTK的FreqDist统计目录高频词遇问题求帮助

Troubleshooting NLTK FreqDist with os.walk for Multi-Directory Text Files

Hey there! Sorry to hear you're hitting roadblocks with your word frequency counting task—let's work through this together.

First, let's start with a working example that uses os.walk and NLTK's FreqDist to scan multiple subdirectories and text files. You can compare this to your code to spot differences:

import os
import nltk
from nltk.tokenize import word_tokenize
from nltk.corpus import stopwords
from nltk.probability import FreqDist

# Download required NLTK resources (run once)
nltk.download('punkt')
nltk.download('stopwords')

def count_high_freq_words(root_dir):
    # Initialize stopwords and frequency distribution
    stop_words = set(stopwords.words('english'))
    all_words = []
    
    # Walk through all directories and files
    for dirpath, _, filenames in os.walk(root_dir):
        for filename in filenames:
            # Only process text files (adjust extension if needed)
            if filename.endswith('.txt'):
                file_path = os.path.join(dirpath, filename)
                try:
                    # Read file content (use appropriate encoding if needed)
                    with open(file_path, 'r', encoding='utf-8') as f:
                        text = f.read()
                    
                    # Tokenize and clean text
                    tokens = word_tokenize(text.lower())
                    # Filter out stopwords, punctuation, and non-alphabetic tokens
                    filtered_tokens = [token for token in tokens if token.isalpha() and token not in stop_words]
                    all_words.extend(filtered_tokens)
                except Exception as e:
                    print(f"Error processing {file_path}: {str(e)}")
    
    # Calculate frequency distribution
    fdist = FreqDist(all_words)
    # Get top 20 high-frequency words
    top_words = fdist.most_common(20)
    
    return top_words

# Replace with your root directory path
root_directory = "/path/to/your/main/directory"
top_freq_words = count_high_freq_words(root_directory)
print("Top 20 High-Frequency Words:")
for word, count in top_freq_words:
    print(f"{word}: {count}")

Now, to zero in on why your code isn't running, could you share a few more details:

  • The exact error message you're seeing (e.g., import failures, file reading errors, tokenization crashes)
  • A snippet of your current code (so we can spot syntax or logic gaps)
  • Quick notes about your text files: do they use a standard encoding like UTF-8? Are there any unusual characters or formatting that might trip up reading/tokenization?

Common pitfalls to double-check on your end:

  • Forgetting to download NLTK resources like punkt or stopwords (this is a super common gotcha!)
  • Mishandling file paths (always use os.path.join() to combine directory and file names)
  • Encoding mismatches when reading files (try adding encoding='utf-8' or encoding='latin-1' to your open() call)
  • Not filtering out non-word tokens (like punctuation or numbers) which can mess up frequency counts
  • Accidentally resetting your frequency distribution instead of accumulating tokens across files

Share those details, and we'll get your code up and running in no time!

内容的提问来源于stack exchange,提问作者NLTK_User

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:26:22