多文件词频统计异常:仅获取最后一个文件词频而非总频次
Fixing Word Frequency Accumulation Across Multiple Files
Got it, I can see exactly what's happening here—you're resetting your word frequency dictionary every time you process a new file, so all the previous counts get overwritten, leaving only the last file's data. Let's fix that by initializing the dictionary once before you start looping through files, then updating it incrementally with each file's word counts.
Here's the revised code with clear explanations:
import csv import glob import re def main(): # Initialize the frequency dictionary ONCE outside the file loop total_word_counts = {} file_list = glob.glob(TARGET_FILES) # Ensure TARGET_FILES is defined, e.g., "*.txt" or "docs/*.md" for file in file_list: with open(file, 'r', encoding='UTF-8', errors='ignore') as f_in: doc = f_in.read() # Update our global count dictionary with words from this file update_counts(doc, total_word_counts) # After processing all files, export the total counts to CSV export_to_csv(total_word_counts, "total_word_frequencies.csv") def update_counts(doc, counts_dict): # Clean and normalize the text: lowercase, remove unwanted characters, split into words cleaned_text = re.sub(r'[^a-zA-Z0-9\s\']', '', doc).lower() # Keeps apostrophes for contractions words = cleaned_text.split() for word in words: # Increment count: use get() to handle words not yet in the dict counts_dict[word] = counts_dict.get(word, 0) + 1 def export_to_csv(counts_dict, output_path): # Write the accumulated counts to a CSV file with open(output_path, 'w', newline='', encoding='UTF-8') as f_out: writer = csv.writer(f_out) writer.writerow(["Word", "Total Frequency"]) # Sort by frequency (highest first) for easier analysis (optional) for word, count in sorted(counts_dict.items(), key=lambda x: x[1], reverse=True): writer.writerow([word, count]) if __name__ == "__main__": main()
Key Fixes & Improvements:
- Persistent Dictionary:
total_word_countsis created once before the file loop, so it retains counts across all files instead of being reset each time. - Incremental Updates: The
update_countsfunction modifies the existing dictionary rather than creating a new one, ensuring we build up total frequencies. - Clean Text Handling: Added basic text normalization (lowercase, retaining apostrophes) to avoid counting "Hello" and "hello" as separate words.
- Modular Code: Split logic into separate functions for readability and easier debugging.
Quick Notes:
- Replace
TARGET_FILESwith your actual file pattern (e.g.,"documents/*.txt"to target all text files in adocumentsfolder). - Adjust the regex in
update_countsif you need to keep or remove specific characters (e.g., remove\'from the regex to strip apostrophes).
Content of the question来源于stack exchange,提问作者john
相关产品推荐
相关产品推荐

