Java统计text2.txt单词在text1.txt出现次数运行结果为0问题排查
代码问题分析及修正方案
核心错误点
- Scanner 指针未重置:
scan1仅在初始化时读取一次text1.txt,第一个单词遍历结束后Scanner指针已经到文件末尾,后续单词遍历的scan1.hasNext()直接返回false,无统计逻辑执行,所以所有结果都是0。 - 计数变量混用:原本用来记录text2单词总数的
wordCount被用来累加单个单词的出现次数,既破坏了待统计单词的遍历边界,也无法得到每个单词的正确计数。 - 匹配逻辑不准确:直接调用
String.contains()会出现子串误匹配,比如文本中的theater会被判定为包含the,不符合独立单词统计的要求。
修正后可运行代码
import java.io.File; import java.util.Scanner; public class WordCountTest { public static void main(String[] args) { try { File file1 = new File("text1.txt"); File file2 = new File("text2.txt"); Scanner scan2 = new Scanner(file2); // 预读取text1全部内容,避免反复读取IO StringBuilder text1Builder = new StringBuilder(); Scanner scan1 = new Scanner(file1); while (scan1.hasNextLine()) { // 前后拼接空格避免行首尾单词匹配异常,不需要大小写不敏感统计可删除toLowerCase() text1Builder.append(" ").append(scan1.nextLine().toLowerCase()).append(" "); } scan1.close(); String text1AllContent = text1Builder.toString(); int targetWordCount = 0; String[] targetWords = new String[100000]; // 读取text2所有待统计单词 while (scan2.hasNextLine()) { String word = scan2.nextLine().trim(); if (!word.isBlank()) { targetWords[targetWordCount] = word.toLowerCase(); targetWordCount++; } } scan2.close(); // 逐个统计每个单词的出现次数 for (int i = 0; i < targetWordCount; i++) { String currentWord = targetWords[i]; int count = 0; int fromIndex = 0; // 前后加空格保证匹配独立单词 String matchStr = " " + currentWord + " "; while ((fromIndex = text1AllContent.indexOf(matchStr, fromIndex)) != -1) { count++; fromIndex += matchStr.length(); } System.out.println(currentWord + " 在 " + file1.getName() + " 中出现次数为: " + count); } } catch (Exception e) { System.out.println("运行出错: \n" + e); } } }
运行结果示例
以你提供的测试文本为例,运行输出如下:
will 在 text1.txt 中出现次数为: 2 dog 在 text1.txt 中出现次数为: 1 the 在 text1.txt 中出现次数为: 3 movie 在 text1.txt 中出现次数为: 1 find 在 text1.txt 中出现次数为: 2
内容的提问来源于stack exchange,提问作者jaramore
相关产品推荐
相关产品推荐

