You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Pymongo加速从字典列表创建MongoDB文档?

加速MongoDB批量文档写入的优化方案

原代码的性能瓶颈

  • 循环中每次执行count_documents和两次update_one,产生大量独立网络请求,通信开销极大
  • 未利用MongoDB的批量操作能力,单条请求的处理效率远低于批量请求

优化方案

方案1:批量预处理+单次批量插入

先批量查询已存在的Id,筛选出不存在的文档并统一添加更新时间后批量插入,大幅减少网络交互次数:

import datetime
from pymongo import MongoClient

client = MongoClient(mongo_connection_string)
db = client.database
coll = db.collection

# 提取所有待处理的Id
all_ids = [record['Id'] for record in dictlist]
# 查询已存在的Id,用集合存储提升判断效率
existing_ids = {doc['Id'] for doc in coll.find({'Id': {'$in': all_ids}}, {'Id': 1})}

# 统一生成更新时间,避免循环内重复计算
time_now = datetime.datetime.utcnow().isoformat()[:-3]
# 筛选出不存在的文档并添加时间字段
new_records = []
for record in dictlist:
    if record['Id'] not in existing_ids:
        record['DocumentUpdatedAt'] = time_now
        new_records.append(record)

# 批量插入新文档
if new_records:
    coll.insert_many(new_records)

方案2:使用bulk_write执行批量更新

如果后续需要更灵活的批量操作(比如支持更新已有文档),可以用bulk_write把所有操作打包为单个请求,同时合并两次update_one为单次操作:

import datetime
from pymongo import MongoClient, UpdateOne

client = MongoClient(mongo_connection_string)
db = client.database
coll = db.collection

time_now = datetime.datetime.utcnow().isoformat()[:-3]
operations = []
for record in dictlist:
    # 合并原记录和更新时间字段,单次完成更新
    update_op = UpdateOne(
        {'Id': record['Id']},
        {'$set': {**record, 'DocumentUpdatedAt': time_now}},
        upsert=False
    )
    operations.append(update_op)

# 批量执行操作,ordered=False允许并行处理,提升速度(无需顺序时推荐使用)
if operations:
    coll.bulk_write(operations, ordered=False)

额外优化建议

  • 给Id字段创建单键索引:coll.create_index('Id'),这会大幅提升查询、判断存在性和更新操作的速度
  • 若无需保证操作顺序,始终使用ordered=False,MongoDB可并行处理批量任务
  • 避免在循环内重复计算相同值(比如time_now),提前计算一次复用

内容的提问来源于stack exchange,提问作者espogian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 22:45:41