为何Pandas允许用np.nan替换0却无法用pd.NA?脚本运行报错
Pandas replace方法使用pd.NA触发递归深度超出错误的问题
问题现象
- 可正常执行的代码:
df = df.replace(0, np.nan) - 异常代码:在Jupyter Notebook中运行正常,但保存为.py文件执行时报错
df = df.replace(0, pd.NA)
复现代码
import numpy as np import pandas as pd df1 = pd.DataFrame({0, 0, 0, 0}) df2 = pd.DataFrame({0, 0, 0, 0}) df1.replace(0, np.nan, inplace=True) df2.replace(0, pd.NA, inplace=True) print(df1) print(df2)
报错信息
File site-packages\pandas\core\dtypes\base.py", line 312, in is_dtype if isinstance(dtype, (ABCSeries, ABCIndex, ABCDataFrame, np.dtype)): File "lib\site-packages\pandas\core\dtypes\generic.py", line 47, in _instancecheck return _check(inst) and not isinstance(inst, type) File "lib\site-packages\pandas\core\dtypes\generic.py", line 41, in _check return getattr(inst, attr, "_typ") in comp RecursionError: maximum recursion depth exceeded while calling a Python object
版本信息
- Pandas版本:1.5.1
- Numpy版本:1.23.4
- Python版本:3.10
原因及解决办法
这是Pandas 1.5.x版本的已知bug,使用pd.NA作为替换值时,会触发类型检查的递归调用循环。Jupyter Notebook因环境初始化逻辑差异暂时规避了该问题,但常规.py脚本会暴露此bug。
解决方式:
- 升级Pandas:更新到1.5.2及以上稳定版本,官方已修复该递归调用问题。
- 临时替代方案:
- 数值型列直接用
np.nan替代pd.NA,二者在数值型数据中行为一致; - 对非数值型列,先转换为可空数据类型再替换:
df2 = df2.astype("Int64") # 转为可空整数类型 df2.replace(0, pd.NA, inplace=True) - 改用
df.where方法实现替换逻辑:df2 = df2.where(df2 != 0, pd.NA)
- 数值型列直接用
内容的提问来源于stack exchange,提问作者Pkmstr
相关产品推荐
相关产品推荐

