You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Java 8 Streams API优化字符串数组的词频统计?

使用Java Streams API优化单词频率统计

需求说明

给定以下Java字符串数组作为输入:

String[] input = new String[] {
       "This is a sample string",
       " string ",                   // 首尾空格会导致拆分问题
       "Another sample string",
       "This is not    a sample string"
};

需要统计每个单词的出现频率,预期输出为:

{a=2, not=1, string=4, This=2, is=2, sample=3, Another=1}

现有粗糙实现

目前已有一段可运行但不够优雅的代码:

// 1. 将String[]拼接成以空格分隔的字符串
String joined = String.join(" ", input);

// 2. 按空格拆分后统计词频
Map<String, Long> output = 
        Arrays
            .stream(joined.split(" "))
            .filter(s -> !s.equals(""))    // 处理拆分产生的空字符串
            .collect(
                Collectors.groupingBy(
                    Function.identity(),
                    Collectors.counting()
                )
            );

System.out.println(output);

优化后的Streams实现

原方案的问题在于:拼接字符串后用split(" ")拆分,会因原数组中的首尾空格、连续空格产生大量空字符串,需要额外过滤,效率和可读性都不佳。下面是更优的实现:

方案一:正则匹配提取单词

import java.util.Arrays;
import java.util.Map;
import java.util.function.Function;
import java.util.regex.Pattern;
import java.util.stream.Collectors;

public class WordFrequency {
    public static void main(String[] args) {
        String[] input = new String[] {
               "This is a sample string",
               " string ",
               "Another sample string",
               "This is not    a sample string"
        };

        // 匹配任意非空白字符序列,直接提取单词
        Pattern wordPattern = Pattern.compile("\\S+");

        Map<String, Long> wordFrequency = Arrays.stream(input)
                .flatMap(str -> wordPattern.matcher(str).results())
                .map(match -> match.group())
                .collect(Collectors.groupingBy(
                        Function.identity(),
                        Collectors.counting()
                ));

        System.out.println(wordFrequency);
    }
}

方案二:正则拆分后扁平化处理

如果偏好拆分的方式,也可以用正则匹配空白字符拆分单个字符串,再合并流统计:

import java.util.Arrays;
import java.util.Map;
import java.util.function.Function;
import java.util.stream.Collectors;

public class WordFrequency {
    public static void main(String[] args) {
        String[] input = new String[] {
               "This is a sample string",
               " string ",
               "Another sample string",
               "This is not    a sample string"
        };

        Map<String, Long> wordFrequency = Arrays.stream(input)
                .map(str -> str.split("\\s+")) // 按任意数量空白字符拆分
                .flatMap(Arrays::stream)
                .filter(s -> !s.isEmpty()) // 处理全空白字符串拆分出的空串
                .collect(Collectors.groupingBy(
                        Function.identity(),
                        Collectors.counting()
                ));

        System.out.println(wordFrequency);
    }
}

优化点说明

  • 精准处理空白字符:用\\S+或\\s+正则,自动跳过首尾空格、连续空格,无需额外过滤大量空字符串
  • 避免冗余操作:直接对单个字符串处理,不需要先拼接成大字符串,减少内存开销
  • 逻辑更连贯:流式操作一气呵成,从提取单词到统计词频的流程清晰可读

内容的提问来源于stack exchange,提问作者Aakash B

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 10:08:12