一维数组高频词查找方法异常:预期返回fat 7却返回There 1求排查
问题:统计数组中最频繁单词结果异常
测试文本:
"There is a very fat cat, the fat is something else. fat fat fat fat. Come on fat is the most used word."
预期返回"fat 7",但实际返回"There 1"。尝试过将track>maxFrequency的判断移到内层for循环外、移出嵌套if结构,结果依旧不对。相关代码如下:
public String findFrequentWord() { String finalP = ""; int maxFrequency = 0; // most used word count int track= 0; //keeping count of how many times a word appears for(int start = 0; start<textList.size();start++){ // picks 1 word at a time for(int next = 0; next<textList.size();next++){ // looks for that word in the entire array before moving to the next word in the list if(textList.get(start)==textList.get(next)){ //when the currently examined word shows up in the list track++; // that words apperence count goes up if(track>maxFrequency){ // before going onto the next word we check if the words apperence is higher than the current max frquency maxFrequency = track; // if it is we update the new maxiumum finalP = textList.get(start); //we update the word displayed the most } //nested if end } // if end } //inner (next) for loop ends track = 0; // reset tracking number for the next word } // outter loop ends return finalP + " " + maxFrequency; // display the most used word and how many times it showed up }
问题分析与修复
核心问题1:String比较误用==而非equals()
Java中==比较的是String对象的内存引用,而非字符串内容。你的textList里的每个String都是独立对象,哪怕内容相同,==也会返回false,导致只有start==next(同一个元素自身比较)时才会触发track++,所以每个单词的track值都只能到1,第一个单词"There"就成了初始最大值,后续无法更新。
核心问题2:统计逻辑时机错误
你在每次track++后就更新最大值,这会导致同一个单词的多次匹配重复更新max,但更关键的是,因为第一个问题的存在,这个逻辑根本没生效。正确做法是等内层循环完全统计完当前单词的总次数后,再和maxFrequency比较,保证统计的是完整次数。
修复后的代码
public String findFrequentWord() { String finalP = ""; int maxFrequency = 0; int track; for(int start = 0; start < textList.size(); start++){ track = 0; // 将track初始化移到内层循环前,逻辑更清晰 for(int next = 0; next < textList.size(); next++){ // 改用equals比较字符串内容 if(textList.get(start).equals(textList.get(next))){ track++; } } // 内层循环结束后,再比较当前单词总次数与最大值 if(track > maxFrequency){ maxFrequency = track; finalP = textList.get(start); } } return finalP + " " + maxFrequency; }
额外优化建议
如果数组规模较大,双层循环的时间复杂度为O(n²),效率较低。可以改用HashMap<String, Integer>统计次数,一次遍历完成统计,时间复杂度降至O(n):
public String findFrequentWord() { Map<String, Integer> countMap = new HashMap<>(); // 遍历统计每个单词的出现次数 for(String word : textList){ countMap.put(word, countMap.getOrDefault(word, 0) + 1); } // 找出次数最多的单词 String finalP = ""; int maxFrequency = 0; for(Map.Entry<String, Integer> entry : countMap.entrySet()){ if(entry.getValue() > maxFrequency){ maxFrequency = entry.getValue(); finalP = entry.getKey(); } } return finalP + " " + maxFrequency; }
内容的提问来源于stack exchange,提问作者N3wb1e
相关产品推荐
相关产品推荐

