You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Java统计文件夹下所有TXT文件中的高频词汇

文件夹全量TXT单词频次统计改造方案

改造后完整代码

import java.io.File;
import java.io.FileNotFoundException;
import java.util.HashMap;
import java.util.Scanner;
import java.util.Map;

public class WordCount {
    // 抽取单个文件统计逻辑,复用全局Map累加所有文件的单词频次
    private static void countWordFromFile(File file, Map<String, Integer> countMap) throws FileNotFoundException {
        // 自动跳过非TXT后缀的文件
        if (!file.getName().toLowerCase().endsWith(".txt")) {
            return;
        }
        Scanner txtFile = new Scanner(file);
        while (txtFile.hasNext()) {
            String word = txtFile.next();
            countMap.put(word, countMap.getOrDefault(word, 0) + 1);
        }
        txtFile.close();
    }

    // 递归遍历文件夹,支持识别子目录内的TXT文件
    private static void traverseFolder(File folder, Map<String, Integer> countMap) throws FileNotFoundException {
        if (!folder.isDirectory()) {
            countWordFromFile(folder, countMap);
            return;
        }
        File[] files = folder.listFiles();
        if (files == null) {
            return;
        }
        for (File file : files) {
            if (file.isDirectory()) {
                traverseFolder(file, countMap);
            } else {
                countWordFromFile(file, countMap);
            }
        }
    }

    public static void main(String[] args) throws FileNotFoundException {
        HashMap<String, Integer> map = new HashMap<>();
        // 此处替换为你需要统计的目标文件夹路径
        File targetFolder = new File("C:\\Users\\Desktop\\testfolder");
        traverseFolder(targetFolder, map);

        for (Map.Entry<String, Integer> entry : map.entrySet()) {
            System.out.println(entry);
        }
    }
}

核心改动点

  • 把单个文件的单词统计逻辑抽为独立公共方法,全局复用同一个Map累加所有文件的统计结果,不会出现多文件统计结果分开的问题
  • 新增文件夹递归遍历逻辑,同时兼容直接传入单个TXT文件的场景,子文件夹内的TXT文件也会被纳入统计
  • 新增TXT后缀校验,自动跳过目录内非TXT格式的文件,避免格式不匹配导致的运行报错
  • 用Map.getOrDefault简化原有频次计数的if-else判断,代码更简洁易读

可选优化

  • 若需要忽略单词大小写统计,可将String word = txtFile.next();修改为String word = txtFile.next().toLowerCase();
  • 若文件夹层级极深担心递归栈溢出,可将递归遍历逻辑替换为栈/队列实现的迭代遍历

内容的提问来源于stack exchange,提问作者Vagelis1234

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 22:36:08