Java 8 Stream两种去重写法差异:为何Set创建次数不同?
按name字段对Employee列表去重的两种写法差异解析
场景描述
想要通过name字段对Employee列表进行去重,编写了两种filter写法,但结果截然不同,以下是具体代码和分析:
第一种写法(去重失效)
public static void main(final String[] args) { final Employee employee = new Employee("test", "123"); final Employee employee1 = new Employee("demo", "3232"); final Employee employee2 = new Employee("test", "323"); final Employee employee3 = new Employee("hello", "123"); final List<Employee> employees = List.of(employee, employee1, employee2, employee3); final List<Employee> collect = employees.stream() .filter(it -> { System.out.println("filter" + it.getName()); return distinctByKey().test(it); }) .collect(Collectors.toList()); System.out.println(collect); System.out.println(seen); } private static Predicate<Employee> distinctByKey() { final Set<String> seen = ConcurrentHashMap.newKeySet(); System.out.println("set"+seen); return employee -> { System.out.println("keyExtractor" + employee.getName()); return seen.add(employee.getName()); }; }
对应输出
filtertest set[] keyExtractortest filterdemo set[] keyExtractordemo filtertest set[] keyExtractortest filterhello set[] keyExtractorhello [Employee{name='test', address='123'}, Employee{name='demo', address='3232'}, Employee{name='test', address='323'}, Employee{name='hello', address='123'}]
可以看到去重未生效,所有元素都被保留。
第二种写法(去重生效)
仅修改filter的调用方式:
final List<Employee> collect = employees.stream() .filter(distinctByKey()) .collect(Collectors.toList());
此时Set仅被创建一次,重复name的元素会被过滤,去重功能正常。
差异原因解析
两种写法的核心区别在于**distinctByKey()方法的调用时机和次数**:
- 第一种写法中,
distinctByKey()写在filter的lambda表达式内部。流中的每个元素进入filter判断时,都会执行一次distinctByKey()方法,而该方法每次执行都会新建一个空的Set<String> seen。这意味着每个元素都在独立的Set里做"是否存在"判断,所有name都是第一次被添加到当前Set,seen.add()永远返回true,所有元素都通过filter,去重自然失效。 - 第二种写法中,
distinctByKey()在创建流的filter环节只调用一次,得到一个绑定了唯一seen集合的Predicate实例。之后流中所有元素的filter判断都会复用这个Predicate,共享同一个seen集合。当遇到重复name时,seen.add()返回false,该元素会被过滤掉,从而实现去重。
简单来说,第一种是每个元素对应一个新Set,第二种是所有元素共享同一个Set,这就是去重效果天差地别的原因。
内容的提问来源于stack exchange,提问作者Remo
相关产品推荐
相关产品推荐

