实现单词统计方法的技术咨询:分割与排序问题(附代码)
Java单词统计问题的解决方案
1. 分割字符串时忽略特殊字符(支持非英文)
原来的split("[ ,.-?]")仅能处理指定的少数特殊字符,无法覆盖所有非字母符号,也不支持非英文Unicode字母。正确的处理方式是利用Unicode字符类匹配所有非字母字符,同时处理分割后产生的空字符串:
- 使用
split("\\P{L}+"):\\P{L}匹配所有非Unicode字母的字符,+表示匹配一个或多个,可一次性分割所有非字母符号(包括中文标点、特殊符号等)。 - 分割后需判断字符串是否为空,避免空字符串被错误统计。
修改mapOfWords方法的分割逻辑:
// 替换原有split与循环逻辑 String[] wordList = line.split("\\P{L}+"); for (String word : wordList) { // 跳过空字符串 if (word.isEmpty()) { continue; } if (map.containsKey(word)) { map.put(word, map.get(word) + 1); } else { map.put(word, 1); } }
2. 按值的数值顺序、再按键的字母顺序排序
HashMap本身是无序的,需将Map的Entry转换为List后自定义排序规则:
- 将
HashMap的entrySet()转为ArrayList,便于排序操作。 - 自定义
Comparator<Map.Entry<String, Integer>>:- 优先比较值(频次),按降序排列(若需升序可反转比较逻辑);
- 当值相同时,按键的自然字母顺序(利用
String.compareTo(),支持Unicode字符排序)排列。
- 调用
Collections.sort()对List执行排序。
修改countWords方法的逻辑:
public String countWords(List<String> lines) { String result = ""; HashMap<String, Integer> wordMap = mapOfWords(lines); // 将Entry转为List以支持排序 List<Map.Entry<String, Integer>> entryList = new ArrayList<>(wordMap.entrySet()); // 自定义排序规则 Collections.sort(entryList, new Comparator<Map.Entry<String, Integer>>() { @Override public int compare(Map.Entry<String, Integer> o1, Map.Entry<String, Integer> o2) { // 先按值降序比较 int valueCompare = o2.getValue().compareTo(o1.getValue()); if (valueCompare != 0) { return valueCompare; } // 值相同时按键的自然顺序升序比较 return o1.getKey().compareTo(o2.getKey()); } }); // 遍历排序后的列表,过滤掉"长度不足4且出现次数少于10"的单词 for (Map.Entry<String, Integer> entry : entryList) { String word = entry.getKey(); int count = entry.getValue(); // 保留:长度>=4 或 次数>=10 的单词(排除长度<4且次数<10的) if (!(word.length() < 4 && count < 10)) { result += word + " - " + count + "\n"; } } // 移除末尾多余的换行符 if (!result.isEmpty()) { result = result.substring(0, result.length() - 1); } return result; }
完整修改后的代码
import java.util.ArrayList; import java.util.Collections; import java.util.Comparator; import java.util.HashMap; import java.util.List; import java.util.Map; public class Words { public String countWords(List<String> lines) { String result = ""; HashMap<String, Integer> wordMap = mapOfWords(lines); List<Map.Entry<String, Integer>> entryList = new ArrayList<>(wordMap.entrySet()); Collections.sort(entryList, new Comparator<Map.Entry<String, Integer>>() { @Override public int compare(Map.Entry<String, Integer> o1, Map.Entry<String, Integer> o2) { int valueCompare = o2.getValue().compareTo(o1.getValue()); if (valueCompare != 0) { return valueCompare; } return o1.getKey().compareTo(o2.getKey()); } }); for (Map.Entry<String, Integer> entry : entryList) { String word = entry.getKey(); int count = entry.getValue(); if (!(word.length() < 4 && count < 10)) { result += word + " - " + count + "\n"; } } if (!result.isEmpty()) { result = result.substring(0, result.length() - 1); } return result; } private HashMap<String, Integer> mapOfWords(List<String> lines) { HashMap<String, Integer> map = new HashMap<>(); for (String line : lines) { // 按所有非Unicode字母分割字符串 String[] wordList = line.split("\\P{L}+"); for (String word : wordList) { if (word.isEmpty()) { continue; } if (map.containsKey(word)) { map.put(word, map.get(word) + 1); } else { map.put(word, 1); } } } return map; } }
关键说明
- 非英文支持:
\\P{L}能匹配所有非Unicode字母,确保中文、日文等非英文语言的字符可被正确分割。 - 过滤逻辑:严格遵循需求,仅排除“长度不足4且出现次数少于10”的单词,其余情况均保留。
- 排序逻辑:先按频次降序排列,频次相同的单词按字母自然顺序升序排列,完全符合需求。
内容的提问来源于stack exchange,提问作者Kere
相关产品推荐
相关产品推荐

