You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Jupyter Notebook中向MongoDB导入CSV文件(百万级数据)

在Jupyter Notebook中将大体积CSV导入MongoDB的两种方案(数据量≥100万条)

针对数据量≥100万条的CSV文件,以下提供两种在Jupyter Notebook中导入MongoDB的实现方案,分别适配不同内存环境:

一、分块导入方案(适合内存有限场景)

当CSV文件体积过大时,分块读取可避免一次性加载全部数据导致的内存过载问题。通过设定chunksize参数拆分数据,分批导入MongoDB,同时统计导入耗时。

from pymongo import MongoClient
import pandas as pd
from time import perf_counter

# 连接MongoDB
client = MongoClient("mongodb://localhost:27017/")
db = client["mflix"]
collection = db["test"]

# 启动计时
t1_start = perf_counter()
# 分块读取CSV,避免内存过载
# 大文件需设置low_memory=False,sep根据文件类型选择(逗号对应csv,制表符对应tsv)
for chunk in pd.read_csv("202201-citibike-tripdata.csv", chunksize=10000, sep=',', low_memory=False):
    chunk.reset_index(inplace=True)
    data_dict = chunk.to_dict("records")
    collection.insert_many(data_dict)
# 停止计时
t1_stop = perf_counter()
# 输出导入耗时
print(f"分块导入总耗时:{t1_stop - t1_start} 秒")

二、非分块导入方案(适合内存充足场景)

若服务器或本地内存足够容纳整个CSV文件,可直接一次性读取全部数据后批量插入,实现逻辑更简洁。

from pymongo import MongoClient
import pandas as pd
from time import perf_counter

# 连接MongoDB
client = MongoClient("mongodb://localhost:27017/")
db = client["mflix"]
collection = db["test"]

# 启动计时
t1_start = perf_counter()
# 一次性读取CSV,大文件需设置low_memory=False
file = pd.read_csv("202201-citibike-tripdata.csv", low_memory=False)
file.reset_index(inplace=True)

data_dict = file.to_dict("records")
collection.insert_many(data_dict)
# 停止计时
t1_stop = perf_counter()
# 输出导入耗时
print(f"非分块导入总耗时:{t1_stop - t1_start} 秒")

内容的提问来源于stack exchange,提问作者anastase

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 09:40:42