如何使用数据类型字典转换Pandas DataFrame列类型并将非法值设为NaN
Pandas 批量转换列数据类型方案(错误值自动转NaN)
实现逻辑
我们可以通过封装通用转换函数,适配类型映射字典中不同的目标类型,所有无法转换的值统一触发coerce规则转成NaN,同时兼容多DataFrame批量调用需求。
完整实现代码
首先导入依赖库:
import pandas as pd import numpy as np from datetime import datetime
通用转换函数
def dtype_convert(df: pd.DataFrame, dtype_dict: dict) -> pd.DataFrame: for col, target_type in dtype_dict.items(): # 跳过当前DataFrame不存在的列 if col not in df.columns: continue # 数值类型(int/float)转换 if target_type in (int, float): df[col] = pd.to_numeric(df[col], errors="coerce") # 可按需指定是否强转目标类型,注意原生int不支持NaN会转成float df[col] = df[col].astype(target_type, errors="ignore") # 日期类型转换 elif target_type == datetime: df[col] = pd.to_datetime(df[col], errors="coerce") # 其他自定义类型转换规则可在此扩展 else: df[col] = df[col].astype(target_type, errors="ignore") return df
调用示例
样例数据与类型字典
# 构造测试DataFrame d = {'col1': ["1", "abc"], 'col2': ["abc", "02-02-2021"]} df = pd.DataFrame(data=d) # 类型映射字典 dtype_dict = { "col1": int, "col2": datetime }
执行转换
df_out = dtype_convert(df, dtype_dict) print(df_out)
输出结果完全符合预期:
col1 col2 0 1.0 NaT 1 NaN 2021-02-02
大规模DataFrame优化建议
- 转换数值类型时可添加
downcast参数缩小内存占用,例如pd.to_numeric(df[col], errors="coerce", downcast="integer") - 若需要保留整数类型同时支持空值,可将类型字典中的
int替换为Pandas可空整数类型pd.Int64Dtype(),避免NaN自动转成float的问题 - 多个DataFrame只需分别传入对应的类型映射字典调用函数即可,无需重复开发转换逻辑
注意:日期转换时如果你的日期格式不是默认兼容格式,可在
pd.to_datetime中添加format参数指定格式,提升转换效率与准确率。
内容的提问来源于stack exchange,提问作者Koot6133
相关产品推荐
相关产品推荐

