Java处理4GB孟加拉语词频文件:合并重复单词并累加对应频率
问题原因
- 核心错误是HashMap的实例化位置放在了while循环内部,每次读取一行就会创建一个全新的空Map,之前所有行的统计数据都会被清空,相当于每次仅处理当前行,自然输出和原文件完全一致。
- 写入文件的逻辑也放在了循环内部,每读一行就立刻写入当前Map的内容,没有等所有行统计完成再统一输出。
- 输出流开启了追加模式,多次运行会导致内容重复叠加到文件尾部。
修正后代码
package data_correction; import java.awt.Toolkit; import java.io.BufferedWriter; import java.io.FileInputStream; import java.io.FileOutputStream; import java.io.OutputStreamWriter; import java.util.HashMap; import java.util.Map; import java.util.Scanner; public class Main { public static void main(String args[]) throws Exception { FileInputStream inputStream = null; Scanner sc = null; String path = "C:\\DATA\\vocab.txt"; // 第二个参数改为false,每次运行覆盖旧的输出文件,避免内容叠加 FileOutputStream fos = new FileOutputStream("C:\\DATA\\output.txt", false); BufferedWriter bufferedWriter = new BufferedWriter(new OutputStreamWriter(fos, "UTF-8")); // HashMap移到循环外部,全程共用一个Map统计所有单词的频率总和 Map<String, Integer> map = new HashMap<>(); try { System.out.println("Started!!"); inputStream = new FileInputStream(path); sc = new Scanner(inputStream, "UTF-8"); while (sc.hasNextLine()) { String line = sc.nextLine().trim(); // 跳过空行避免数组越界 if (line.isEmpty()) continue; String[] arr = line.split("="); // 校验格式,避免不符合key=value的行报错 if (arr.length < 2) continue; String key = arr[0]; int count = Integer.parseInt(arr[1].trim()); map.put(key, map.getOrDefault(key, 0) + count); } // 所有行统计完成后,统一遍历Map写入结果 for (Map.Entry<String, Integer> each : map.entrySet()) { bufferedWriter.write(each.getKey() + "=" + each.getValue() + "\n"); } bufferedWriter.close(); if (sc.ioException() != null) { throw sc.ioException(); } } finally { if (inputStream != null) { inputStream.close(); } if (sc != null) { sc.close(); } } System.out.print("FINISH"); Toolkit.getDefaultToolkit().beep(); } }
补充说明
如果文件对应的单词总量极大,内存无法容纳所有键值对,可以改用外部排序+归并的方式处理,但常规自然语言的词汇总量通常在百万级以内,上述代码完全可以正常运行。
内容的提问来源于stack exchange,提问作者user16930946
相关产品推荐
相关产品推荐

