You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Dataprep API遇RecursionError及缺失值填充代码优化咨询

问题描述

在Google Colab中使用Dataprep API时,处理含144个特征的未清洗数据一切正常,但将数据精简至20个特征并完成缺失值清洗后,调用create_report()函数触发RecursionError。经排查确定是缺失值填充代码存在问题,相关代码如下:

# 需用中位数填充缺失值的列
median_fill = ['emp_length','annual_inc','open_acc','pub_rec','open_acc','revol_util','total_acc']
for med in median_fill:
  df[med].fillna(df[med].median,inplace=True)

执行相关代码及报错信息:

df.isna().sum()
# 输出各列缺失值均为0
sys.setrecursionlimit(15000)
from dataprep.eda import create_report, plot, plot_correlation
create_report(df)

报错信息:

RecursionError: maximum recursion depth exceeded
问题根源

代码中df[med].median是方法对象引用,而非调用方法后得到的中位数数值。填充时将方法对象存入了DataFrame的单元格中,后续Dataprep的create_report()在解析这些单元格内容时,会陷入无限递归解析方法对象的逻辑,最终触发递归深度超限错误。

优化后的代码

将df[med].median改为df[med].median(),调用方法获取实际的中位数数值再填充:

# 需用中位数填充缺失值的列
median_fill = ['emp_length','annual_inc','open_acc','pub_rec','open_acc','revol_util','total_acc']
for med in median_fill:
    # 调用median()方法获取中位数数值
    median_val = df[med].median()
    df[med].fillna(median_val, inplace=True)

或者更高效的向量化写法,避免循环:

median_fill = ['emp_length','annual_inc','open_acc','pub_rec','open_acc','revol_util','total_acc']
df[median_fill] = df[median_fill].fillna(df[median_fill].median())
补充说明
  1. 向量化写法效率更高,尤其适合处理较大数据集
  2. 填充后建议再次检查df.dtypes,确保目标列仍为数值类型,避免因存入方法对象导致类型异常

内容的提问来源于stack exchange,提问作者Gregory Morris

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 11:31:13