You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

遍历文件统计词频函数结果重复,请求代码问题排查

问题排查与代码修正

核心问题分析

你遇到的结果重复问题,主要由以下几个代码错误导致:

  • CSV保存逻辑错误:循环内每次都将累计的全量统计结果写入当前文件对应的路径(且路径拼接错误),最终所有生成的CSV都会包含所有文件的统计数据,看起来结果完全相同。
  • 文件筛选逻辑错误:filter(files, '*.txt')用法不符合Python语法,无法正确筛选出txt文件,可能导致处理的文件列表异常。
  • 缩进错误:统计、数据生成等代码错误缩进在with open块内,虽然不直接引发重复,但容易导致逻辑混乱。
  • 词汇拼写错误:部分统计词汇拼写不全(如cryptocurrenc应为cryptocurrency),会导致统计结果失真。

修正后的代码

import os
from os.path import join
import pandas as pd
from os import walk

def process_files(target_dir):
    # 正确筛选所有子文件夹中的txt文件
    filenames = []
    for root, _, files in walk(target_dir):
        # 用endswith筛选txt文件,替代错误的filter用法
        txt_files = [join(root, file) for file in files if file.endswith('.txt')]
        filenames.extend(txt_files)
    
    df = pd.DataFrame()
    # 定义要统计的词汇,避免重复变量声明
    target_words = {
        'Virtual Currency': 'virtual currency',
        'Bitcoin': 'bitcoin',
        'Cryptocurrency': 'cryptocurrency',
        'Digital Currency': 'digital currency',
        'Litecoin': 'litecoin',
        'Dogecoin': 'dogecoin',
        'Ethereum': 'ethereum'
    }

    for filename in filenames:
        # 初始化统计结果
        word_counts = {'File Name': filename}
        
        with open(filename, 'r', encoding='utf-8') as f:
            contents = f.read().lower()
        
        # 遍历词汇统计,简化代码
        for col_name, word in target_words.items():
            word_counts[col_name] = contents.count(word)
        
        # 将当前文件的统计结果添加到DataFrame
        df = pd.concat([df, pd.DataFrame([word_counts])], ignore_index=True)
    
    # 将所有文件的统计汇总结果保存到目标文件夹的汇总CSV中
    df.to_csv(join(target_dir, 'word_count_summary.csv'), index=False)
    # 如果需要为每个文件单独生成统计CSV,取消下面的注释
    # for _, row in df.iterrows():
    #     single_df = pd.DataFrame([row])
    #     # 生成单独的文件名,比如在原文件名后加"_count.csv"
    #     single_csv_path = row['File Name'].replace('.txt', '_count.csv')
    #     single_df.to_csv(single_csv_path, index=False)
    
    return df

result = process_files(r'C:\test\QTR2')

关键修复点说明

  1. 文件筛选:用file.endswith('.txt')替代错误的filter用法,确保只处理txt文件。
  2. CSV保存逻辑:将汇总结果的保存移到循环外,生成一个统一的汇总CSV;如果需要每个文件单独的统计文件,可启用注释部分的代码,为每个文件生成独立的统计CSV(避免路径拼接错误)。
  3. 代码简化:用字典存储目标词汇,避免重复声明变量,提升代码可维护性。
  4. 拼写修正:修正了词汇拼写错误(如etherrum改为ethereum,补全currency的后缀),确保统计准确。
  5. 缩进修正:将统计、数据生成等代码移出with open块,保证文件读取完成后再处理数据。

内容的提问来源于stack exchange,提问作者Ciercy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 22:47:51