Mongoose批量更新存量数据适配新库结构及代码优化
MongoDB存量数据结构迁移性能优化方案
现有实现的性能瓶颈根因
当前代码执行效率低是以下几个硬伤导致的:
- 全量
find({})一次性拉取所有记录到应用内存,数据量级较大时会直接导致内存占用飙升,查询阶段耗时极长,甚至触发OOM forEach传入async回调不会被外层等待,相当于瞬间发起和总数据量同量级的save()请求,直接打满数据库连接池,触发数据库流控- 单条记录单独调用
save(),每一条更新都要走一次网络往返、数据库写日志流程,IO开销被放大数百倍 - 用
JSON.parse(JSON.stringify())做深拷贝属于完全多余的操作,白白消耗CPU性能
方案1:原生聚合管道更新(性能最优,推荐)
如果使用MongoDB 4.2及以上版本,直接用聚合管道完成全量更新,所有转换逻辑在数据库层执行,不需要把数据拉到应用服务,性能是所有方案里最高的,百万级数据通常几分钟就能跑完。
async function convertOldToNew() { await this.progressModel.updateMany( // 过滤条件:仅处理未完成迁移的记录,支持中断重跑,避免重复执行 { $or: [ { "answers.name": { $type: "string" } }, { "habitos.residuos.name": { $exists: true } }, { "habitos.consumo.name": { $exists: true } }, { "habitos.transporte.name": { $exists: true } } ] }, [ // 转换answers数组下的name字段为{es: 原值}格式 { $set: { answers: { $map: { input: { $ifNull: ["$answers", []] }, in: { $mergeObjects: [ "$$this", { name: { es: "$$this.name" } } ] } } } } }, // 移除habitos下三个数组元素内的冗余name、how、why字段 { $set: { "habitos.residuos": { $map: { input: { $ifNull: ["$habitos.residuos", []] }, in: { $arrayToObject: { $filter: { input: { $objectToArray: "$$this" }, cond: { $not: { $in: ["$$this.k", ["name", "how", "why"]] } } } } } } }, "habitos.consumo": { $map: { input: { $ifNull: ["$habitos.consumo", []] }, in: { $arrayToObject: { $filter: { input: { $objectToArray: "$$this" }, cond: { $not: { $in: ["$$this.k", ["name", "how", "why"]] } } } } } } }, "habitos.transporte": { $map: { input: { $ifNull: ["$habitos.transporte", []] }, in: { $arrayToObject: { $filter: { input: { $objectToArray: "$$this" }, cond: { $not: { $in: ["$$this.k", ["name", "how", "why"]] } } } } } } } } } ] ) }
该方案优势:
- 无应用层数据传输、内存开销,所有计算在数据库侧完成
- 仅需一次数据库请求即可完成全量更新,IO开销降到最低
- 内置过滤条件,脚本中途中断可以直接重跑,不会重复处理已迁移数据
方案2:分块查询+BulkWrite批量写入(兼容低版本MongoDB)
如果使用的MongoDB版本低于4.2,不支持聚合管道更新,可以用分块游标拉取+批量写入的方案,兼顾性能和稳定性,千万级数据也能平稳执行。
async function convertOldToNew() { const BATCH_SIZE = 1000; // 每批处理条数,可根据数据库性能调整到1000-5000 let lastId = null; const redundantFields = ['name', 'how', 'why']; const habitosKeys = ['residuos', 'consumo', 'transporte']; while (true) { // 按_id分块查询,每次仅拉取一批数据,避免全量加载内存 const query = lastId ? { _id: { $gt: lastId } } : {}; const batch = await this.progressModel .find(query) .select({ habitos: 1, answers: 1 }) .sort({ _id: 1 }) .limit(BATCH_SIZE) .lean(); // 直接返回普通JS对象,省去Mongoose文档实例化开销,性能提升3-5倍,无需额外深拷贝 if (batch.length === 0) break; const bulkOperations = []; for (const progress of batch) { // 处理answers字段转换 if (progress.answers?.length) { progress.answers.forEach(element => { if (typeof element.name === 'string') { element.name = { es: element.name }; } }); } // 处理habitos下冗余字段删除 if (progress.habitos) { habitosKeys.forEach(key => { if (progress.habitos[key]?.length) { progress.habitos[key] = progress.habitos[key].map(element => { redundantFields.forEach(field => delete element[field]); return element; }) } }) } // 加入批量更新队列 bulkOperations.push({ updateOne: { filter: { _id: progress._id }, update: { $set: { answers: progress.answers, habitos: progress.habitos } } } }) } // 批量提交更新,每批仅一次网络请求 await this.progressModel.bulkWrite(bulkOperations, { ordered: false }); lastId = batch[batch.length - 1]._id; console.log(`已完成${batch.length}条记录迁移,当前最后处理ID:${lastId}`); } }
该方案优化点:
- 用
lean()查询替代默认Mongoose文档查询,省去模型实例化、数据监听的额外开销,也不需要用JSON序列化做深拷贝 - 按_id有序分块拉取,每次仅加载少量数据到内存,不会出现内存溢出
- 用
bulkWrite批量提交更新,相比单条save()网络IO开销降低99%以上 - 配置
ordered: false允许MongoDB乱序执行写入,进一步提升批量写入性能
迁移注意事项
- 执行迁移前务必备份全量数据,先在测试环境验证转换逻辑正确性
- 批量大小不要设置超过5000,避免单次请求的BSON大小超过MongoDB 16MB的限制
- 尽量在业务低峰期执行迁移,避免占用过多数据库CPU、IO资源影响线上业务
- 脚本支持断点续跑,中途中断后直接重新执行即可,会自动跳过已完成迁移的记录
内容的提问来源于stack exchange,提问作者Abdul Rehman Dar
相关产品推荐
相关产品推荐

