如何统计文件中单词出现次数?HashMap与ArrayList方案遇阻
嘿,我明白你现在的困扰了——想用HashMap统计单词频次却卡在了“索引访问”的问题上,转而考虑用两个ArrayList分别存单词和对应次数对吧?咱们一步步来拆解问题,先聊聊HashMap的使用误区,再看看你想的ArrayList方案怎么落地,最后给你几个更高效的实现思路。
先纠正HashMap的“索引访问”误区
其实HashMap本来就不是靠索引访问的,它的核心优势是通过单词(key)直接定位出现次数(value),这反而比用ArrayList更高效。你觉得它无法正常工作,大概率是没掌握正确的使用姿势。
举个Java代码的实际例子,正确的HashMap统计逻辑应该是这样的:
import java.util.HashMap; import java.util.Scanner; import java.io.File; import java.io.FileNotFoundException; public class WordCounter { public static void main(String[] args) { HashMap<String, Integer> wordCountMap = new HashMap<>(); try (Scanner scanner = new Scanner(new File("your-target-file.txt"))) { while (scanner.hasNext()) { // 统一转小写+去除标点,避免"Hello"和"hello"被当成不同单词 String word = scanner.next().toLowerCase().replaceAll("[^a-zA-Z]", ""); if (!word.isEmpty()) { // 用getOrDefault简化判断:如果单词不存在,默认次数为0,加1后存入 wordCountMap.put(word, wordCountMap.getOrDefault(word, 0) + 1); } } } catch (FileNotFoundException e) { e.printStackTrace(); } // 如果需要按插入顺序遍历(类似ArrayList的索引顺序),换成LinkedHashMap即可 // 遍历输出所有单词和次数 for (String word : wordCountMap.keySet()) { System.out.println(word + ": " + wordCountMap.get(word)); } } }
如果需要保持单词的插入顺序(比如和文件里出现的顺序一致),只要把HashMap换成LinkedHashMap就行——它既保留了HashMap的O(1)高效查找,又能按插入顺序遍历元素。
再说说你想用的两个ArrayList方案
如果坚持要用两个ArrayList来实现,思路完全可行,只是效率会稍低一些,适合小文件场景。具体逻辑是:
- 用
ArrayList<String> words存储所有不重复的单词; - 用
ArrayList<Integer> counts存储对应索引位置单词的出现次数; - 读取每个单词时,先检查
words中是否已存在该单词:- 存在的话,找到对应索引,把
counts里的数值加1; - 不存在的话,把单词加入
words,同时往counts里添加初始值1。
- 存在的话,找到对应索引,把
代码示例如下:
import java.util.ArrayList; import java.util.Scanner; import java.io.File; import java.io.FileNotFoundException; public class WordCounterWithLists { public static void main(String[] args) { ArrayList<String> words = new ArrayList<>(); ArrayList<Integer> counts = new ArrayList<>(); try (Scanner scanner = new Scanner(new File("your-target-file.txt"))) { while (scanner.hasNext()) { String word = scanner.next().toLowerCase().replaceAll("[^a-zA-Z]", ""); if (!word.isEmpty()) { int index = words.indexOf(word); if (index != -1) { // 单词已存在,对应次数加1 counts.set(index, counts.get(index) + 1); } else { // 单词不存在,新增到两个列表中 words.add(word); counts.add(1); } } } } catch (FileNotFoundException e) { e.printStackTrace(); } // 按索引遍历输出 for (int i = 0; i < words.size(); i++) { System.out.println(words.get(i) + ": " + counts.get(i)); } } }
⚠️ 注意:indexOf是线性查找,当文件里的单词数量很多时,效率会比HashMap低很多(HashMap查找是O(1),ArrayList是O(n)),所以大文件场景更推荐用HashMap系列。
额外推荐:Java 8+的Stream API简化实现
如果你的项目用的是Java 8及以上版本,还可以用Stream API写更简洁的代码,一行核心逻辑就能完成统计:
import java.io.IOException; import java.nio.file.Files; import java.nio.file.Paths; import java.util.Map; import java.util.stream.Collectors; public class WordCounterStream { public static void main(String[] args) { try { Map<String, Long> wordCount = Files.lines(Paths.get("your-target-file.txt")) .flatMap(line -> java.util.Arrays.stream(line.split("\\W+"))) .map(String::toLowerCase) .filter(word -> !word.isEmpty()) .collect(Collectors.groupingBy(word -> word, Collectors.counting())); wordCount.forEach((word, count) -> System.out.println(word + ": " + count)); } catch (IOException e) { e.printStackTrace(); } } }
这里用Files.lines读取文件行,拆分单词后直接分组统计,代码更优雅易读。
内容的提问来源于stack exchange,提问作者M.Ahmed
相关产品推荐
相关产品推荐

