You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Jupyter中用DuckDB加载大Parquet表时内核崩溃求助

Jupyter Notebook中DuckDB加载超大Parquet文件时内核崩溃的原因与解决方法

问题场景

尝试在Jupyter Notebook通过Python调用DuckDB加载超大Parquet文件为表,代码如下:

authorDf = conn.execute("""
CREATE TABLE author AS SELECT * FROM read_parquet('/home/data/author.parquet');
""")

执行后Jupyter会话崩溃,收到错误:

Canceled future for execute_request message before replies were done
The Kernel crashed while executing code in the the current cell or a previous cell. Please review the code in the cell(s) to identify a possible cause of the failure. Click here for more info.

可能原因

  • 内存资源耗尽:DuckDB默认会尝试将全量数据加载到内存中,超大Parquet文件直接触发内存溢出,导致内核崩溃。
  • Jupyter内核资源限制:Jupyter进程本身的内存配额不足,无法支撑大文件的加载操作。
  • Parquet文件异常:文件损坏、schema定义错误等问题,也可能导致加载过程中出现不可预期的崩溃(这类情况通常伴随更明确的错误提示,但也存在无提示崩溃的可能)。

解决方法

1. 配置DuckDB内存管理与磁盘溢出

通过设置DuckDB的参数,限制内存使用并启用磁盘临时存储,避免内存溢出:

# 设置内存限制(根据你的机器内存调整,比如16GB)
conn.execute("PRAGMA memory_limit='16GB'")
# 指定临时文件目录(需确保目录所在磁盘有足够空间)
conn.execute("PRAGMA temp_directory='/home/data/tmp'")
# 再执行建表操作
authorDf = conn.execute("""
CREATE TABLE author AS SELECT * FROM read_parquet('/home/data/author.parquet');
""")

2. 避免全量加载,按需读取

不要使用SELECT *加载所有列,只选择需要的字段;或者添加过滤条件减少数据量:

-- 仅加载需要的列
CREATE TABLE author AS SELECT id, name, email FROM read_parquet('/home/data/author.parquet');

-- 或添加过滤条件
CREATE TABLE author AS SELECT * FROM read_parquet('/home/data/author.parquet') WHERE join_year >= 2020;

3. 调整Jupyter内核资源限制

启动Jupyter时增加内存缓冲区大小,例如设置为16GB:

jupyter notebook --NotebookApp.max_buffer_size=16106127360

也可以修改Jupyter的配置文件(jupyter_notebook_config.py),永久设置资源限制。

4. 检查Parquet文件完整性

先尝试读取少量数据验证文件是否正常:

# 读取前10条数据测试
test_df = conn.execute("SELECT * FROM read_parquet('/home/data/author.parquet') LIMIT 10").df()
print(test_df)

如果这一步也崩溃,说明Parquet文件可能损坏,需要重新生成或修复文件。

内容的提问来源于stack exchange,提问作者PabloM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 05:07:27