pandas加载数据集指定dtype时如何直接实现错误容错处理
问题解答
pd.read_csv() 没有内置和pd.to_numeric(..., errors='coerce')直接对应的全局参数,仅通过dtype传类型字典时,只要某列存在无法转换为目标类型的异常值,读取流程会直接报错,不会自动将异常值容错为空值。可以通过以下两种方式实现你要的效果:
方案1:读取后逐列做容错转换(最推荐,逻辑可控)
这种方式不会增加读取阶段的性能开销,转换规则灵活可自定义:
- 第一步先将所有列按字符串类型读入,避免读取阶段的自动类型转换提前吞掉异常格式
- 第二步遍历定义的类型映射字典,针对不同类型调用带
errors='coerce'的转换方法,转换失败的值会自动置为空值
示例代码:
import pandas as pd import numpy as np # 自定义列类型映射 columns2type = { "column1": str, "column2": pd.Int64Dtype(), # 存在空值的整数列请用pandas可空整数类型,不要用原生int "column3": float } # 先全量按字符串读入 df = pd.read_csv("path/file", dtype=str) # 逐列做带容错的类型转换 for col, target_type in columns2type.items(): if pd.api.types.is_numeric_dtype(target_type): df[col] = pd.to_numeric(df[col], errors="coerce").astype(target_type) elif pd.api.types.is_datetime64_dtype(target_type): df[col] = pd.to_datetime(df[col], errors="coerce") else: df[col] = df[col].astype(target_type)
方案2:通过converters参数在读取阶段完成转换
如果不想在读取完成后做二次处理,可以给需要容错的列传入自定义转换函数,在读取过程中就捕获类型转换异常,返回空值:
import pandas as pd import numpy as np columns2type = { "column1": str, "column2": pd.Int64Dtype(), "column3": float } # 构造带容错的数值转换函数 def make_coerce_converter(target_dtype): def _converter(value): try: return target_dtype(value) except (ValueError, TypeError): return np.nan return _converter # 分离需要转换器的列和直接指定类型的列 converters = {} direct_dtype = {} for col, dtype in columns2type.items(): if pd.api.types.is_numeric_dtype(dtype): converters[col] = make_coerce_converter(dtype) else: direct_dtype[col] = dtype df = pd.read_csv("path/file", converters=converters, dtype=direct_dtype)
注意:
converters参数的优先级高于dtype,配置了转换器的列不要重复在dtype里传数值类型,避免冲突。
内容的提问来源于stack exchange,提问作者Manolo Dominguez Becerra
相关产品推荐
相关产品推荐

