如何解决使用Polars时出现的'Canceled future for execute_request message'错误
Polars处理35GB文件内核崩溃的修复方案
问题描述
使用Polars 0.16.6处理35GB文件时,执行以下代码触发内核崩溃:
df.unique(subset=['id']).select(pl.count()).collect(streaming=True).item()
Jupyter提示错误:
Canceled future for execute_request message before replies were done
The Kernel crashed while executing code in the current cell or a previous cell. Please review the code in the cell(s) to identify a possible cause of the failure. Click here for more info. View Jupyter log for further details.
查看日志发现核心错误是内存分配失败:memory allocation of 301430844 bytes failed,并非延迟求值导致。
修复方法
- 升级Polars版本
0.16.6版本对streaming模式下的unique操作优化不足,处理大基数唯一ID时会占用大量内存。升级到0.17+的最新稳定版,官方优化了该场景的内存使用:
pip install --upgrade polars
- 分块处理唯一ID
如果暂时无法升级,可通过分块读取文件,先统计每个块的唯一ID,再合并去重计数,避免一次性加载所有数据:
import polars as pl # 分块读取并收集每个块的唯一ID chunk_unique_list = [] for chunk in pl.scan_parquet("your_large_file.parquet").select("id").streaming().iter_batches(): chunk_unique_list.append(chunk.unique(subset=['id'])) # 合并所有块的唯一ID,最终统计总数 total_unique_count = pl.concat(chunk_unique_list).unique(subset=['id']).select(pl.count()).item() print(total_unique_count)
- 使用近似计数(非精确场景)
如果不需要精确的唯一值数量,可改用approx_n_unique大幅降低内存消耗:
pl.scan_parquet("your_large_file.parquet").select(pl.col("id").approx_n_unique()).collect(streaming=True).item()
- 释放系统内存
关闭其他占用内存的程序,为Polars进程预留足够内存;Windows用户可检查是否开启了内存限制,尝试调整Python进程的可用内存。
内容的提问来源于stack exchange,提问作者roei shlezinger
相关产品推荐
相关产品推荐

