You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多文件词频统计异常:仅获取最后一个文件词频而非总频次

Fixing Word Frequency Accumulation Across Multiple Files

Got it, I can see exactly what's happening here—you're resetting your word frequency dictionary every time you process a new file, so all the previous counts get overwritten, leaving only the last file's data. Let's fix that by initializing the dictionary once before you start looping through files, then updating it incrementally with each file's word counts.

Here's the revised code with clear explanations:

import csv
import glob
import re

def main():
    # Initialize the frequency dictionary ONCE outside the file loop
    total_word_counts = {}
    file_list = glob.glob(TARGET_FILES)  # Ensure TARGET_FILES is defined, e.g., "*.txt" or "docs/*.md"
    
    for file in file_list:
        with open(file, 'r', encoding='UTF-8', errors='ignore') as f_in:
            doc = f_in.read()
            # Update our global count dictionary with words from this file
            update_counts(doc, total_word_counts)
    
    # After processing all files, export the total counts to CSV
    export_to_csv(total_word_counts, "total_word_frequencies.csv")

def update_counts(doc, counts_dict):
    # Clean and normalize the text: lowercase, remove unwanted characters, split into words
    cleaned_text = re.sub(r'[^a-zA-Z0-9\s\']', '', doc).lower()  # Keeps apostrophes for contractions
    words = cleaned_text.split()
    
    for word in words:
        # Increment count: use get() to handle words not yet in the dict
        counts_dict[word] = counts_dict.get(word, 0) + 1

def export_to_csv(counts_dict, output_path):
    # Write the accumulated counts to a CSV file
    with open(output_path, 'w', newline='', encoding='UTF-8') as f_out:
        writer = csv.writer(f_out)
        writer.writerow(["Word", "Total Frequency"])
        # Sort by frequency (highest first) for easier analysis (optional)
        for word, count in sorted(counts_dict.items(), key=lambda x: x[1], reverse=True):
            writer.writerow([word, count])

if __name__ == "__main__":
    main()

Key Fixes & Improvements:

  • Persistent Dictionary: total_word_counts is created once before the file loop, so it retains counts across all files instead of being reset each time.
  • Incremental Updates: The update_counts function modifies the existing dictionary rather than creating a new one, ensuring we build up total frequencies.
  • Clean Text Handling: Added basic text normalization (lowercase, retaining apostrophes) to avoid counting "Hello" and "hello" as separate words.
  • Modular Code: Split logic into separate functions for readability and easier debugging.

Quick Notes:

  • Replace TARGET_FILES with your actual file pattern (e.g., "documents/*.txt" to target all text files in a documents folder).
  • Adjust the regex in update_counts if you need to keep or remove specific characters (e.g., remove \' from the regex to strip apostrophes).

Content of the question来源于stack exchange,提问作者john

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:12:07