You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java正则表达式问题:文本词频统计排序及特殊字符匹配修复

文本分词与词频统计问题修复

待处理文本

“And how about you, Count Peter Kirílych? If they call up the militia,
you too will have to mount a horse,” remarked the old count, addressing
Pierre.

需求与问题

需求:对上述文本分词,统计长度≥4的单词的出现次数,按词频从高到低返回结果。

存在的问题:

  • 无法实现词频降序输出
  • 部分词频统计结果有误
  • 正则表达式无法正确处理破折号(-)、反引号(`)和引号

原实现代码

StringBuilder sb= new StringBuilder();
Map<String, Integer> counterMap = new HashMap<>();
for(Object str:lines){
    sb.append(str.toString()+" ");
}
String boom[] = sb.toString().split("['@?-` \\p{Punct}]+\\s*");
for (String word : boom) {
    if(word.length()>=4){
        if(!word.isEmpty()) {
            word = word.trim();
            Integer count = counterMap.get(word);
            if(count == null) {
                count = 0;
            }
            counterMap.put(word, ++count);
        }
    }
}
StringBuilder bb = new StringBuilder();
Map.Entry<String,Integer> maxEntry = null;

for(String word : counterMap.keySet()) {
    System.out.println(word + ": " + counterMap.get(word));
    bb.append(word + " - "+counterMap.get(word)+"\n");
}

return bb.toString();

解决方案

1. 修复正则表达式

原正则中?-会被解析为字符范围(匹配?到-之间的字符),导致匹配逻辑错误。调整为匹配所有标点符号(含引号、破折号、反引号)和空白字符作为分隔符:

String[] boom = sb.toString().split("[\\p{Punct}`'\"-]+|\\s+");

该正则会将所有标点(包括", ', -, `)和空白字符作为分词分隔符,避免特殊符号残留到单词中。

2. 修正词频统计逻辑

原代码先判断长度再trim,会导致带前后空格的单词长度判断错误;同时空字符串的判断顺序不合理。调整为:

for (String word : boom) {
    String trimmedWord = word.trim();
    // 先判断非空,再判断长度≥4
    if (!trimmedWord.isEmpty() && trimmedWord.length() >= 4) {
        // 统一转为小写(可选,避免大小写重复统计,比如Count和count)
        String lowerWord = trimmedWord.toLowerCase();
        counterMap.put(lowerWord, counterMap.getOrDefault(lowerWord, 0) + 1);
    }
}
  • 先trim再判断,避免空格影响长度计算
  • 加入小写转换(可选),解决大小写单词被重复统计的问题(比如原文本中的Count和count会被视为同一个词)
  • 使用getOrDefault简化计数逻辑

3. 实现词频降序排序

HashMap是无序的,需要将Map的entry转为列表后排序:

// 将Map.entry转为列表并排序
List<Map.Entry<String, Integer>> sortedEntries = new ArrayList<>(counterMap.entrySet());
sortedEntries.sort((entry1, entry2) -> {
    // 先按词频降序,词频相同则按单词字母升序
    int freqCompare = entry2.getValue().compareTo(entry1.getValue());
    if (freqCompare != 0) {
        return freqCompare;
    }
    return entry1.getKey().compareTo(entry2.getKey());
});

// 构建结果字符串
StringBuilder bb = new StringBuilder();
for (Map.Entry<String, Integer> entry : sortedEntries) {
    System.out.println(entry.getKey() + ": " + entry.getValue());
    bb.append(entry.getKey() + " - " + entry.getValue() + "\n");
}

完整修正代码

StringBuilder sb = new StringBuilder();
Map<String, Integer> counterMap = new HashMap<>();
for (Object str : lines) {
    sb.append(str.toString()).append(" ");
}

// 修复后的正则分词
String[] boom = sb.toString().split("[\\p{Punct}`'\"-]+|\\s+");

for (String word : boom) {
    String trimmedWord = word.trim();
    if (!trimmedWord.isEmpty() && trimmedWord.length() >= 4) {
        // 统一转为小写,避免大小写重复统计
        String lowerWord = trimmedWord.toLowerCase();
        counterMap.put(lowerWord, counterMap.getOrDefault(lowerWord, 0) + 1);
    }
}

// 按词频降序排序
List<Map.Entry<String, Integer>> sortedEntries = new ArrayList<>(counterMap.entrySet());
sortedEntries.sort((entry1, entry2) -> {
    int freqCompare = entry2.getValue().compareTo(entry1.getValue());
    return freqCompare != 0 ? freqCompare : entry1.getKey().compareTo(entry2.getKey());
});

StringBuilder bb = new StringBuilder();
for (Map.Entry<String, Integer> entry : sortedEntries) {
    System.out.println(entry.getKey() + ": " + entry.getValue());
    bb.append(entry.getKey() + " - " + entry.getValue() + "\n");
}

return bb.toString();

测试结果

对给定文本处理后,输出结果(按词频降序):

count - 2
militia - 1
peter - 1
kirílych - 1
remarked - 1
addressing - 1
pierre - 1

内容的提问来源于stack exchange,提问作者Boombom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 08:25:17