如何从MongoDB聚合查询中获取含重复项的全部numberId信息
解决方法
你当前的聚合查询返回去重结果的核心原因是使用了group("numBerId")阶段——这个操作会将所有拥有相同numBerId的文档合并为一个分组,自然会去除重复项。要获取包含重复项的所有numBerId字段,只需移除group阶段,改用project提取目标字段即可:
修改后的聚合代码如下:
Aggregation agg = TypedAggregation.newAggregation( TypedAggregation.match(Criteria.where("numBerId").regex("^" + numBerId, "i") .andOperator(Criteria.where("numBerId").ne(""))), TypedAggregation.project("numBerId"), // 仅保留numBerId字段,替代原group阶段 TypedAggregation.sort(Direction.ASC, "numBerId"), // 直接按numBerId字段排序 TypedAggregation.limit(20000) ); Document rawResults = mongo.aggregate(agg, collectionName(), Document.class).getRawResults(); return rawResults.getList("results", Document.class) .stream() .map(d -> (String) d.get("numBerId")) // 从numBerId字段取值,而非原分组的_id .collect(Collectors.toList());
关键修改说明:
- 移除
group阶段:删除这个分组操作,就能保留所有匹配的原始文档,自然包含重复的numBerId。 - 添加
project阶段:用于筛选出我们需要的numBerId字段,减少不必要的数据传输(若不需要精简数据,也可省略此阶段,但后续仍需从完整文档中提取numBerId)。 - 调整排序与取值逻辑:原代码的排序字段
_id是分组后的标识,现在直接使用numBerId排序;结果提取时也改为从numBerId字段取值,而非分组id。
内容的提问来源于stack exchange,提问作者Андрей Андрей
相关产品推荐
相关产品推荐

