You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java Stream API处理重复templateId:基于source差异保留最大runCount

解决方案

要实现你需要的逻辑,核心是按templateId分组后针对每组做两种情况判断:

  • 组内所有source值一致:保留全部记录
  • 组内存在不同source值:仅保留runCount最大的那条记录

修改后的代码如下:

List<Entity> autoGens = repository.findAllByRemoved(false);

List<Entity> autoGensResult = autoGens.stream()
    // 按templateId分组,得到每个模板ID对应的所有实体列表
    .collect(Collectors.groupingBy(Entity::getTemplateId))
    .values()
    .stream()
    .flatMap(group -> {
        // 判断当前组内所有source是否完全相同
        boolean allSourcesSame = group.stream()
            .map(Entity::getSource)
            .distinct()
            .count() == 1;
        
        if (allSourcesSame) {
            // source值统一,保留组内所有记录
            return group.stream();
        } else {
            // source存在差异,筛选出runCount最大的那条记录
            return group.stream()
                .max(Comparator.comparing(Entity::getRunCount))
                .stream(); // 把Optional转为Stream,适配flatMap的统一处理
        }
    })
    .collect(Collectors.toList());

代码说明

  1. 分组处理:先用groupingBy将实体按templateId拆分,得到每个模板对应的全量实体集合。
  2. source一致性校验:通过distinct().count() == 1快速判断组内source是否无差异。
  3. 分支逻辑:
    • 若source完全一致,直接将组内所有实体流入结果集
    • 若source有差异,找出组内runCount最大的实体,转为Stream后合并到结果流
  4. 结果收集:用flatMap合并所有分支的流,最终收集成目标列表。

内容的提问来源于stack exchange,提问作者Mikhail

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 15:36:25