You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python在MongoDB中获取两个数组的交集大小

嘿,这个需求我之前在项目里碰到过类似的——处理MongoDB里大数组的交集计数,得选对方法才不会搞崩数据库。下面给你两种实用方案,根据你的数据规模挑就行:

方案一:聚合管道直接计算(适合中小规模数组)

如果你的followers数组元素不算特别多(比如几十万以内),用这个方法最直接,一步到位:

// 替换成你的集合名称
db.your_collection.aggregate([
  // 先筛选出要对比的两个账户文档
  { $match: {
      screen_name: { $in: ['BBC', 'TheEconomist'] }
    }
  },
  // 把两个账户的followers数组收集到一个列表里
  { $group: {
      _id: null,
      followerLists: { $push: "$followers" }
    }
  },
  // 计算交集并返回长度
  { $project: {
      commonFollowersCount: {
        $size: {
          $setIntersection: ["$followerLists.0", "$followerLists.1"]
        }
      },
      _id: 0
    }
  }
])

优缺点:代码简洁易懂,上手快;但如果followers是百万级别的超大数组,这个方法会把整个数组加载到内存,容易引发性能瓶颈甚至内存溢出。

方案二:索引优化+聚合(适合大规模数组)

如果你的followers数组真的有几百万元素,那得靠索引来提速,避免全数组加载。

首先给followers字段建个单字段索引:

db.your_collection.createIndex({ followers: 1 })

然后用聚合管道关联两个文档,MongoDB会利用索引优化交集计算:

db.your_collection.aggregate([
  // 第一步:获取第一个账户(比如BBC)的followers数组
  { $match: { screen_name: "BBC" } },
  { $project: { followers: 1, _id: 0 } },
  // 第二步:关联到第二个目标账户(TheEconomist)的文档
  { $lookup: {
      from: "your_collection",
      pipeline: [
        { $match: { screen_name: "TheEconomist" } },
        { $project: { followers: 1, _id: 0 } }
      ],
      as: "targetAccount"
    }
  },
  // 第三步:计算两个数组的交集大小
  { $project: {
      commonFollowersCount: {
        $size: {
          $setIntersection: ["$followers", "$targetAccount.0.followers"]
        }
      },
      _id: 0
    }
  }
])

额外提示:

  • 如果followers数组里有重复的ID,记得先去重,不然交集计数会包含重复项,只需把$setIntersection里的参数改成{$addToSet: "$followers"}和{$addToSet: "$targetAccount.0.followers"}即可。
  • 要是数组规模达到千万级别,建议换个数据结构:把followers拆成单独的集合,每条记录存{ account_id: "xxx", follower_id: "yyy" },然后用count结合$in或者聚合统计共同关注者,这种方式在分片集群下性能会好很多。

内容的提问来源于stack exchange,提问作者pattpattpatt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:41:39