使用os.walk与NLTK的FreqDist统计目录高频词遇问题求帮助
Troubleshooting NLTK FreqDist with os.walk for Multi-Directory Text Files
Hey there! Sorry to hear you're hitting roadblocks with your word frequency counting task—let's work through this together.
First, let's start with a working example that uses os.walk and NLTK's FreqDist to scan multiple subdirectories and text files. You can compare this to your code to spot differences:
import os import nltk from nltk.tokenize import word_tokenize from nltk.corpus import stopwords from nltk.probability import FreqDist # Download required NLTK resources (run once) nltk.download('punkt') nltk.download('stopwords') def count_high_freq_words(root_dir): # Initialize stopwords and frequency distribution stop_words = set(stopwords.words('english')) all_words = [] # Walk through all directories and files for dirpath, _, filenames in os.walk(root_dir): for filename in filenames: # Only process text files (adjust extension if needed) if filename.endswith('.txt'): file_path = os.path.join(dirpath, filename) try: # Read file content (use appropriate encoding if needed) with open(file_path, 'r', encoding='utf-8') as f: text = f.read() # Tokenize and clean text tokens = word_tokenize(text.lower()) # Filter out stopwords, punctuation, and non-alphabetic tokens filtered_tokens = [token for token in tokens if token.isalpha() and token not in stop_words] all_words.extend(filtered_tokens) except Exception as e: print(f"Error processing {file_path}: {str(e)}") # Calculate frequency distribution fdist = FreqDist(all_words) # Get top 20 high-frequency words top_words = fdist.most_common(20) return top_words # Replace with your root directory path root_directory = "/path/to/your/main/directory" top_freq_words = count_high_freq_words(root_directory) print("Top 20 High-Frequency Words:") for word, count in top_freq_words: print(f"{word}: {count}")
Now, to zero in on why your code isn't running, could you share a few more details:
- The exact error message you're seeing (e.g., import failures, file reading errors, tokenization crashes)
- A snippet of your current code (so we can spot syntax or logic gaps)
- Quick notes about your text files: do they use a standard encoding like UTF-8? Are there any unusual characters or formatting that might trip up reading/tokenization?
Common pitfalls to double-check on your end:
- Forgetting to download NLTK resources like
punktorstopwords(this is a super common gotcha!) - Mishandling file paths (always use
os.path.join()to combine directory and file names) - Encoding mismatches when reading files (try adding
encoding='utf-8'orencoding='latin-1'to youropen()call) - Not filtering out non-word tokens (like punctuation or numbers) which can mess up frequency counts
- Accidentally resetting your frequency distribution instead of accumulating tokens across files
Share those details, and we'll get your code up and running in no time!
内容的提问来源于stack exchange,提问作者NLTK_User
相关产品推荐
相关产品推荐

