MongoDB中如何按expId使用distinct去重查询浏览体验数据
Hey there! 作为MongoDB新手碰到重复数据的问题太正常啦,我来帮你搞定按expId去重的需求~
首先得明确你想要的结果类型:是只需要去重后的expId列表,还是要保留完整的文档(每个expId只留一条)?我分两种情况给你解决方案:
1. 只需要获取去重后的expId数组
如果你的需求只是拿到所有不重复的expId,那用MongoDB的distinct()方法最直接。假设你的原查询代码是类似这样的:
// 原查询(返回带重复expId的完整文档) const views = await db.collection('Views').find({ userId: req.userId }).toArray();
只需要改成下面这样,就能直接得到去重后的expId数组:
const uniqueExpIds = await db.collection('Views').distinct('expId', { userId: req.userId });
返回结果会是类似:["exp001", "exp002", "exp003"]
2. 需要保留完整文档,每个expId只留一条
如果要返回完整的文档数据,同时保证每个expId唯一,distinct()就不够用了(它只能返回单个字段的去重结果),这时候得用聚合管道来实现。
方案:按expId分组,保留每组的第一条/最新文档
比如你想保留每个expId对应的最新一条数据(假设你的文档有createdAt时间字段),可以这么写:
const uniqueViews = await db.collection('Views').aggregate([ // 第一步:过滤当前用户的所有数据 { $match: { userId: req.userId } }, // 第二步:按时间降序排序,确保最新的文档排在每组最前面 { $sort: { createdAt: -1 } }, // 第三步:按expId分组,取每组的第一个文档(也就是最新的那条) { $group: { _id: '$expId', // 分组的键是expId doc: { $first: '$$ROOT' } // $$ROOT代表整个文档,$first取每组第一个 } }, // 第四步:把分组后的doc替换成根文档,让结果和find返回的格式一致 { $replaceRoot: { newRoot: '$doc' } } ]).toArray();
如果你用的是Mongoose(Node.js的MongoDB ODM)
写法基本一致,只是调用方式略有不同:
// Mongoose中用distinct const uniqueExpIds = await View.distinct('expId', { userId: req.userId }); // Mongoose中用聚合管道 const uniqueViews = await View.aggregate([ { $match: { userId: req.userId } }, { $sort: { createdAt: -1 } }, { $group: { _id: '$expId', doc: { $first: '$$ROOT' } } }, { $replaceRoot: { newRoot: '$doc' } } ]);
举个例子对比:
原返回结果(带重复expId):
[ { "_id": "doc1", "userId": "user123", "expId": "exp001", "content": "浏览记录1", "createdAt": "2024-01-01" }, { "_id": "doc2", "userId": "user123", "expId": "exp001", "content": "浏览记录2", "createdAt": "2024-01-02" }, { "_id": "doc3", "userId": "user123", "expId": "exp002", "content": "浏览记录3", "createdAt": "2024-01-01" } ]
用聚合管道后的返回结果(expId唯一,保留最新的文档):
[ { "_id": "doc2", "userId": "user123", "expId": "exp001", "content": "浏览记录2", "createdAt": "2024-01-02" }, { "_id": "doc3", "userId": "user123", "expId": "exp002", "content": "浏览记录3", "createdAt": "2024-01-01" } ]
如果你的文档没有时间字段,也可以用$last或者直接按默认顺序取每组的第一条,根据你的需求调整就行~
内容的提问来源于stack exchange,提问作者dev dev
相关产品推荐
相关产品推荐

