You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从MongoDB聚合查询中获取含重复项的全部numberId信息

解决方法

你当前的聚合查询返回去重结果的核心原因是使用了group("numBerId")阶段——这个操作会将所有拥有相同numBerId的文档合并为一个分组,自然会去除重复项。要获取包含重复项的所有numBerId字段,只需移除group阶段,改用project提取目标字段即可:

修改后的聚合代码如下:

Aggregation agg = TypedAggregation.newAggregation(
        TypedAggregation.match(Criteria.where("numBerId").regex("^" + numBerId, "i")
                .andOperator(Criteria.where("numBerId").ne(""))),
        TypedAggregation.project("numBerId"), // 仅保留numBerId字段,替代原group阶段
        TypedAggregation.sort(Direction.ASC, "numBerId"), // 直接按numBerId字段排序
        TypedAggregation.limit(20000)
);

Document rawResults = mongo.aggregate(agg, collectionName(), Document.class).getRawResults();
return rawResults.getList("results", Document.class)
        .stream()
        .map(d -> (String) d.get("numBerId")) // 从numBerId字段取值,而非原分组的_id
        .collect(Collectors.toList());

关键修改说明:

  • 移除group阶段:删除这个分组操作,就能保留所有匹配的原始文档,自然包含重复的numBerId。
  • 添加project阶段:用于筛选出我们需要的numBerId字段,减少不必要的数据传输(若不需要精简数据,也可省略此阶段,但后续仍需从完整文档中提取numBerId)。
  • 调整排序与取值逻辑:原代码的排序字段_id是分组后的标识,现在直接使用numBerId排序;结果提取时也改为从numBerId字段取值,而非分组id。

内容的提问来源于stack exchange,提问作者Андрей Андрей

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 11:35:13