You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决使用Polars时出现的'Canceled future for execute_request message'错误

Polars处理35GB文件内核崩溃的修复方案

问题描述

使用Polars 0.16.6处理35GB文件时,执行以下代码触发内核崩溃:

df.unique(subset=['id']).select(pl.count()).collect(streaming=True).item()

Jupyter提示错误:

Canceled future for execute_request message before replies were done
The Kernel crashed while executing code in the current cell or a previous cell. Please review the code in the cell(s) to identify a possible cause of the failure. Click here for more info. View Jupyter log for further details.

查看日志发现核心错误是内存分配失败:memory allocation of 301430844 bytes failed,并非延迟求值导致。

修复方法

  • 升级Polars版本
    0.16.6版本对streaming模式下的unique操作优化不足,处理大基数唯一ID时会占用大量内存。升级到0.17+的最新稳定版,官方优化了该场景的内存使用:
pip install --upgrade polars
  • 分块处理唯一ID
    如果暂时无法升级,可通过分块读取文件,先统计每个块的唯一ID,再合并去重计数,避免一次性加载所有数据:
import polars as pl

# 分块读取并收集每个块的唯一ID
chunk_unique_list = []
for chunk in pl.scan_parquet("your_large_file.parquet").select("id").streaming().iter_batches():
    chunk_unique_list.append(chunk.unique(subset=['id']))

# 合并所有块的唯一ID,最终统计总数
total_unique_count = pl.concat(chunk_unique_list).unique(subset=['id']).select(pl.count()).item()
print(total_unique_count)
  • 使用近似计数(非精确场景)
    如果不需要精确的唯一值数量,可改用approx_n_unique大幅降低内存消耗:
pl.scan_parquet("your_large_file.parquet").select(pl.col("id").approx_n_unique()).collect(streaming=True).item()
  • 释放系统内存
    关闭其他占用内存的程序,为Polars进程预留足够内存;Windows用户可检查是否开启了内存限制,尝试调整Python进程的可用内存。

内容的提问来源于stack exchange,提问作者roei shlezinger

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 12:22:31