如何使用Java 8 Streams API优化字符串数组的词频统计?
使用Java Streams API优化单词频率统计
需求说明
给定以下Java字符串数组作为输入:
String[] input = new String[] { "This is a sample string", " string ", // 首尾空格会导致拆分问题 "Another sample string", "This is not a sample string" };
需要统计每个单词的出现频率,预期输出为:
{a=2, not=1, string=4, This=2, is=2, sample=3, Another=1}
现有粗糙实现
目前已有一段可运行但不够优雅的代码:
// 1. 将String[]拼接成以空格分隔的字符串 String joined = String.join(" ", input); // 2. 按空格拆分后统计词频 Map<String, Long> output = Arrays .stream(joined.split(" ")) .filter(s -> !s.equals("")) // 处理拆分产生的空字符串 .collect( Collectors.groupingBy( Function.identity(), Collectors.counting() ) ); System.out.println(output);
优化后的Streams实现
原方案的问题在于:拼接字符串后用split(" ")拆分,会因原数组中的首尾空格、连续空格产生大量空字符串,需要额外过滤,效率和可读性都不佳。下面是更优的实现:
方案一:正则匹配提取单词
import java.util.Arrays; import java.util.Map; import java.util.function.Function; import java.util.regex.Pattern; import java.util.stream.Collectors; public class WordFrequency { public static void main(String[] args) { String[] input = new String[] { "This is a sample string", " string ", "Another sample string", "This is not a sample string" }; // 匹配任意非空白字符序列,直接提取单词 Pattern wordPattern = Pattern.compile("\\S+"); Map<String, Long> wordFrequency = Arrays.stream(input) .flatMap(str -> wordPattern.matcher(str).results()) .map(match -> match.group()) .collect(Collectors.groupingBy( Function.identity(), Collectors.counting() )); System.out.println(wordFrequency); } }
方案二:正则拆分后扁平化处理
如果偏好拆分的方式,也可以用正则匹配空白字符拆分单个字符串,再合并流统计:
import java.util.Arrays; import java.util.Map; import java.util.function.Function; import java.util.stream.Collectors; public class WordFrequency { public static void main(String[] args) { String[] input = new String[] { "This is a sample string", " string ", "Another sample string", "This is not a sample string" }; Map<String, Long> wordFrequency = Arrays.stream(input) .map(str -> str.split("\\s+")) // 按任意数量空白字符拆分 .flatMap(Arrays::stream) .filter(s -> !s.isEmpty()) // 处理全空白字符串拆分出的空串 .collect(Collectors.groupingBy( Function.identity(), Collectors.counting() )); System.out.println(wordFrequency); } }
优化点说明
- 精准处理空白字符:用
\\S+或\\s+正则,自动跳过首尾空格、连续空格,无需额外过滤大量空字符串 - 避免冗余操作:直接对单个字符串处理,不需要先拼接成大字符串,减少内存开销
- 逻辑更连贯:流式操作一气呵成,从提取单词到统计词频的流程清晰可读
内容的提问来源于stack exchange,提问作者Aakash B
相关产品推荐
相关产品推荐

