You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MongoDB 4.0(DocumentDB)按用户查找首个无间隙连续key序列的高效方案

MongoDB 4.0(DocumentDB)高效查找用户首个无间隙连续key的文档

针对百万级数据量的场景,要高效实现按用户提取首个无间隙连续key的文档,核心是利用索引优化排序与分组,结合聚合管道完成序列分析。以下是具体实现方案:

第一步:创建高效索引

首先创建复合索引,确保排序和分组操作能利用索引加速,避免全表扫描和内存排序:

db.collection.createIndex({ user: 1, key: 1 })

第二步:聚合管道实现逻辑

通过聚合管道完成排序、分组、连续序列分析和文档筛选,具体代码如下:

db.collection.aggregate([
  // 利用索引快速排序,避免内存排序开销
  { $sort: { user: 1, key: 1 } },
  // 按用户分组,收集排序后的文档和key序列
  {
    $group: {
      _id: "$user",
      sortedDocs: { $push: "$$ROOT" },
      keySequence: { $push: "$key" }
    }
  },
  // 计算首个无间隙连续序列的最大key
  {
    $project: {
      sortedDocs: 1,
      maxContinuousKey: {
        $reduce: {
          input: "$keySequence",
          initialValue: {
            currentExpected: { $arrayElemAt: ["$keySequence", 0] },
            maxKey: { $arrayElemAt: ["$keySequence", 0] },
            isContinuous: true
          },
          in: {
            $cond: {
              if: { 
                $and: [
                  "$$value.isContinuous",
                  { $ne: ["$$this", { $add: ["$$value.currentExpected", 1] }] },
                  { $ne: ["$$this", { $arrayElemAt: ["$keySequence", 0] }] }
                ] 
              },
              then: {
                currentExpected: "$$value.currentExpected",
                maxKey: "$$value.maxKey",
                isContinuous: false
              },
              else: {
                $cond: {
                  if: "$$value.isContinuous",
                  then: {
                    currentExpected: "$$this",
                    maxKey: "$$this",
                    isContinuous: true
                  },
                  else: "$$value"
                }
              }
            }
          }
        }
      }
    }
  },
  // 筛选出连续序列内的文档
  {
    $project: {
      user: "$_id",
      continuousDocs: {
        $filter: {
          input: "$sortedDocs",
          cond: { $lte: ["$$this.key", "$maxContinuousKey.maxKey"] }
        }
      },
      _id: 0
    }
  },
  // 可选:将文档数组展开为单条文档返回
  { $unwind: "$continuousDocs" },
  // 可选:投影出需要的字段
  {
    $project: {
      key: "$continuousDocs.key",
      user: "$user",
      comment: "$continuousDocs.comment"
    }
  }
])

逻辑说明

  1. $sort阶段:依赖提前创建的复合索引,MongoDB直接使用索引完成排序,无需将百万级数据加载到内存,性能最优。
  2. $group阶段:按用户分组,将每个用户的文档按key顺序存入数组,同时收集key序列,输入已排序的前提下分组效率极高。
  3. $reduce阶段:遍历key序列,跟踪连续状态:
    • 以第一个key为初始起点,标记为连续状态。
    • 后续每个key如果等于预期的下一个值(当前预期值+1),则更新预期值和连续序列的最大key;一旦出现间隙,终止连续状态。
  4. $filter阶段:从排序后的文档数组中筛选出key不超过连续序列最大key的文档,得到目标结果。

边界情况处理

  • 若用户只有单条文档:直接返回该文档。
  • 若用户的key序列完全无间隙:返回该用户的所有文档。

内容的提问来源于stack exchange,提问作者user13981607

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 16:40:26