You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

实现单词统计方法的技术咨询:分割与排序问题(附代码)

Java单词统计问题的解决方案

1. 分割字符串时忽略特殊字符(支持非英文)

原来的split("[ ,.-?]")仅能处理指定的少数特殊字符,无法覆盖所有非字母符号,也不支持非英文Unicode字母。正确的处理方式是利用Unicode字符类匹配所有非字母字符,同时处理分割后产生的空字符串:

  • 使用split("\\P{L}+"):\\P{L}匹配所有非Unicode字母的字符,+表示匹配一个或多个,可一次性分割所有非字母符号(包括中文标点、特殊符号等)。
  • 分割后需判断字符串是否为空,避免空字符串被错误统计。

修改mapOfWords方法的分割逻辑:

// 替换原有split与循环逻辑
String[] wordList = line.split("\\P{L}+");
for (String word : wordList) {
    // 跳过空字符串
    if (word.isEmpty()) {
        continue;
    }
    if (map.containsKey(word)) {
        map.put(word, map.get(word) + 1);
    } else {
        map.put(word, 1);
    }
}

2. 按值的数值顺序、再按键的字母顺序排序

HashMap本身是无序的,需将Map的Entry转换为List后自定义排序规则:

  1. 将HashMap的entrySet()转为ArrayList,便于排序操作。
  2. 自定义Comparator<Map.Entry<String, Integer>>:
    • 优先比较值(频次),按降序排列(若需升序可反转比较逻辑);
    • 当值相同时,按键的自然字母顺序(利用String.compareTo(),支持Unicode字符排序)排列。
  3. 调用Collections.sort()对List执行排序。

修改countWords方法的逻辑:

public String countWords(List<String> lines) {
    String result = "";
    HashMap<String, Integer> wordMap = mapOfWords(lines);
    
    // 将Entry转为List以支持排序
    List<Map.Entry<String, Integer>> entryList = new ArrayList<>(wordMap.entrySet());
    
    // 自定义排序规则
    Collections.sort(entryList, new Comparator<Map.Entry<String, Integer>>() {
        @Override
        public int compare(Map.Entry<String, Integer> o1, Map.Entry<String, Integer> o2) {
            // 先按值降序比较
            int valueCompare = o2.getValue().compareTo(o1.getValue());
            if (valueCompare != 0) {
                return valueCompare;
            }
            // 值相同时按键的自然顺序升序比较
            return o1.getKey().compareTo(o2.getKey());
        }
    });
    
    // 遍历排序后的列表,过滤掉"长度不足4且出现次数少于10"的单词
    for (Map.Entry<String, Integer> entry : entryList) {
        String word = entry.getKey();
        int count = entry.getValue();
        // 保留:长度>=4 或 次数>=10 的单词(排除长度<4且次数<10的)
        if (!(word.length() < 4 && count < 10)) {
            result += word + " - " + count + "\n";
        }
    }
    
    // 移除末尾多余的换行符
    if (!result.isEmpty()) {
        result = result.substring(0, result.length() - 1);
    }
    return result;
}

完整修改后的代码

import java.util.ArrayList;
import java.util.Collections;
import java.util.Comparator;
import java.util.HashMap;
import java.util.List;
import java.util.Map;

public class Words {
    public String countWords(List<String> lines) {
        String result = "";
        HashMap<String, Integer> wordMap = mapOfWords(lines);
        
        List<Map.Entry<String, Integer>> entryList = new ArrayList<>(wordMap.entrySet());
        
        Collections.sort(entryList, new Comparator<Map.Entry<String, Integer>>() {
            @Override
            public int compare(Map.Entry<String, Integer> o1, Map.Entry<String, Integer> o2) {
                int valueCompare = o2.getValue().compareTo(o1.getValue());
                if (valueCompare != 0) {
                    return valueCompare;
                }
                return o1.getKey().compareTo(o2.getKey());
            }
        });
        
        for (Map.Entry<String, Integer> entry : entryList) {
            String word = entry.getKey();
            int count = entry.getValue();
            if (!(word.length() < 4 && count < 10)) {
                result += word + " - " + count + "\n";
            }
        }
        
        if (!result.isEmpty()) {
            result = result.substring(0, result.length() - 1);
        }
        return result;
    }

    private HashMap<String, Integer> mapOfWords(List<String> lines) {
        HashMap<String, Integer> map = new HashMap<>();
        for (String line : lines) {
            // 按所有非Unicode字母分割字符串
            String[] wordList = line.split("\\P{L}+");
            for (String word : wordList) {
                if (word.isEmpty()) {
                    continue;
                }
                if (map.containsKey(word)) {
                    map.put(word, map.get(word) + 1);
                } else {
                    map.put(word, 1);
                }
            }
        }
        return map;
    }
}

关键说明

  • 非英文支持:\\P{L}能匹配所有非Unicode字母,确保中文、日文等非英文语言的字符可被正确分割。
  • 过滤逻辑:严格遵循需求,仅排除“长度不足4且出现次数少于10”的单词,其余情况均保留。
  • 排序逻辑:先按频次降序排列,频次相同的单词按字母自然顺序升序排列,完全符合需求。

内容的提问来源于stack exchange,提问作者Kere

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 20:51:37