Java正则表达式问题:文本词频统计排序及特殊字符匹配修复
文本分词与词频统计问题修复
待处理文本
“And how about you, Count Peter Kirílych? If they call up the militia,
you too will have to mount a horse,” remarked the old count, addressing
Pierre.
需求与问题
需求:对上述文本分词,统计长度≥4的单词的出现次数,按词频从高到低返回结果。
存在的问题:
- 无法实现词频降序输出
- 部分词频统计结果有误
- 正则表达式无法正确处理破折号(-)、反引号(`)和引号
原实现代码
StringBuilder sb= new StringBuilder(); Map<String, Integer> counterMap = new HashMap<>(); for(Object str:lines){ sb.append(str.toString()+" "); } String boom[] = sb.toString().split("['@?-` \\p{Punct}]+\\s*"); for (String word : boom) { if(word.length()>=4){ if(!word.isEmpty()) { word = word.trim(); Integer count = counterMap.get(word); if(count == null) { count = 0; } counterMap.put(word, ++count); } } } StringBuilder bb = new StringBuilder(); Map.Entry<String,Integer> maxEntry = null; for(String word : counterMap.keySet()) { System.out.println(word + ": " + counterMap.get(word)); bb.append(word + " - "+counterMap.get(word)+"\n"); } return bb.toString();
解决方案
1. 修复正则表达式
原正则中?-会被解析为字符范围(匹配?到-之间的字符),导致匹配逻辑错误。调整为匹配所有标点符号(含引号、破折号、反引号)和空白字符作为分隔符:
String[] boom = sb.toString().split("[\\p{Punct}`'\"-]+|\\s+");
该正则会将所有标点(包括", ', -, `)和空白字符作为分词分隔符,避免特殊符号残留到单词中。
2. 修正词频统计逻辑
原代码先判断长度再trim,会导致带前后空格的单词长度判断错误;同时空字符串的判断顺序不合理。调整为:
for (String word : boom) { String trimmedWord = word.trim(); // 先判断非空,再判断长度≥4 if (!trimmedWord.isEmpty() && trimmedWord.length() >= 4) { // 统一转为小写(可选,避免大小写重复统计,比如Count和count) String lowerWord = trimmedWord.toLowerCase(); counterMap.put(lowerWord, counterMap.getOrDefault(lowerWord, 0) + 1); } }
- 先trim再判断,避免空格影响长度计算
- 加入小写转换(可选),解决大小写单词被重复统计的问题(比如原文本中的
Count和count会被视为同一个词) - 使用
getOrDefault简化计数逻辑
3. 实现词频降序排序
HashMap是无序的,需要将Map的entry转为列表后排序:
// 将Map.entry转为列表并排序 List<Map.Entry<String, Integer>> sortedEntries = new ArrayList<>(counterMap.entrySet()); sortedEntries.sort((entry1, entry2) -> { // 先按词频降序,词频相同则按单词字母升序 int freqCompare = entry2.getValue().compareTo(entry1.getValue()); if (freqCompare != 0) { return freqCompare; } return entry1.getKey().compareTo(entry2.getKey()); }); // 构建结果字符串 StringBuilder bb = new StringBuilder(); for (Map.Entry<String, Integer> entry : sortedEntries) { System.out.println(entry.getKey() + ": " + entry.getValue()); bb.append(entry.getKey() + " - " + entry.getValue() + "\n"); }
完整修正代码
StringBuilder sb = new StringBuilder(); Map<String, Integer> counterMap = new HashMap<>(); for (Object str : lines) { sb.append(str.toString()).append(" "); } // 修复后的正则分词 String[] boom = sb.toString().split("[\\p{Punct}`'\"-]+|\\s+"); for (String word : boom) { String trimmedWord = word.trim(); if (!trimmedWord.isEmpty() && trimmedWord.length() >= 4) { // 统一转为小写,避免大小写重复统计 String lowerWord = trimmedWord.toLowerCase(); counterMap.put(lowerWord, counterMap.getOrDefault(lowerWord, 0) + 1); } } // 按词频降序排序 List<Map.Entry<String, Integer>> sortedEntries = new ArrayList<>(counterMap.entrySet()); sortedEntries.sort((entry1, entry2) -> { int freqCompare = entry2.getValue().compareTo(entry1.getValue()); return freqCompare != 0 ? freqCompare : entry1.getKey().compareTo(entry2.getKey()); }); StringBuilder bb = new StringBuilder(); for (Map.Entry<String, Integer> entry : sortedEntries) { System.out.println(entry.getKey() + ": " + entry.getValue()); bb.append(entry.getKey() + " - " + entry.getValue() + "\n"); } return bb.toString();
测试结果
对给定文本处理后,输出结果(按词频降序):
count - 2 militia - 1 peter - 1 kirílych - 1 remarked - 1 addressing - 1 pierre - 1
内容的提问来源于stack exchange,提问作者Boombom
相关产品推荐
相关产品推荐

