如何用Python在MongoDB中获取两个数组的交集大小
嘿,这个需求我之前在项目里碰到过类似的——处理MongoDB里大数组的交集计数,得选对方法才不会搞崩数据库。下面给你两种实用方案,根据你的数据规模挑就行:
方案一:聚合管道直接计算(适合中小规模数组)
如果你的followers数组元素不算特别多(比如几十万以内),用这个方法最直接,一步到位:
// 替换成你的集合名称 db.your_collection.aggregate([ // 先筛选出要对比的两个账户文档 { $match: { screen_name: { $in: ['BBC', 'TheEconomist'] } } }, // 把两个账户的followers数组收集到一个列表里 { $group: { _id: null, followerLists: { $push: "$followers" } } }, // 计算交集并返回长度 { $project: { commonFollowersCount: { $size: { $setIntersection: ["$followerLists.0", "$followerLists.1"] } }, _id: 0 } } ])
优缺点:代码简洁易懂,上手快;但如果followers是百万级别的超大数组,这个方法会把整个数组加载到内存,容易引发性能瓶颈甚至内存溢出。
方案二:索引优化+聚合(适合大规模数组)
如果你的followers数组真的有几百万元素,那得靠索引来提速,避免全数组加载。
首先给followers字段建个单字段索引:
db.your_collection.createIndex({ followers: 1 })
然后用聚合管道关联两个文档,MongoDB会利用索引优化交集计算:
db.your_collection.aggregate([ // 第一步:获取第一个账户(比如BBC)的followers数组 { $match: { screen_name: "BBC" } }, { $project: { followers: 1, _id: 0 } }, // 第二步:关联到第二个目标账户(TheEconomist)的文档 { $lookup: { from: "your_collection", pipeline: [ { $match: { screen_name: "TheEconomist" } }, { $project: { followers: 1, _id: 0 } } ], as: "targetAccount" } }, // 第三步:计算两个数组的交集大小 { $project: { commonFollowersCount: { $size: { $setIntersection: ["$followers", "$targetAccount.0.followers"] } }, _id: 0 } } ])
额外提示:
- 如果
followers数组里有重复的ID,记得先去重,不然交集计数会包含重复项,只需把$setIntersection里的参数改成{$addToSet: "$followers"}和{$addToSet: "$targetAccount.0.followers"}即可。 - 要是数组规模达到千万级别,建议换个数据结构:把
followers拆成单独的集合,每条记录存{ account_id: "xxx", follower_id: "yyy" },然后用count结合$in或者聚合统计共同关注者,这种方式在分片集群下性能会好很多。
内容的提问来源于stack exchange,提问作者pattpattpatt
相关产品推荐
相关产品推荐

