如何通过字典转换Python DataFrame中的数据类型?
用字典批量转换DataFrame数据类型的实现方法
核心思路
先基于你提供的类型映射字典,生成DataFrame各列对应的目标类型映射关系,再通过astype()方法批量转换;针对时间类型可单独处理,避免字符串转datetime64时出现解析失败问题。
实现步骤及代码
- 准备类型映射字典(你提供的)
type_conversion = { "bigint": "int64", "boolean": "bool", "character varying": "str", "double precision": "float64", "integer": "int32", "numeric": "float64", "timestamp without time zone": "datetime64" }
- 生成列与目标类型的映射
首先要明确DataFrame每列的当前数据类型名称(需与type_conversion的键匹配,比如从数据库读取的原始类型)。假设你已获取列与当前类型的对应关系(如col_types字典,键为列名,值为当前类型),生成目标映射:
# 仅保留可转换列的目标类型映射 target_type_map = { col: type_conversion[ctype] for col, ctype in col_types.items() if ctype in type_conversion }
如果当前类型是从DataFrame的dtypes获取,可先转换为字符串匹配:
col_types = {col: str(df[col].dtype) for col in df.columns} target_type_map = { col: type_conversion[ctype] for col, ctype in col_types.items() if ctype in type_conversion }
- 批量转换数据类型
- 普通类型直接用
astype()转换:
df = df.astype(target_type_map)
- 时间类型单独处理(适配字符串格式的时间数据):
如果原始时间列是字符串格式,先通过pd.to_datetime()解析,再设置类型:
# 筛选出时间类型列 time_cols = [col for col, ctype in col_types.items() if ctype == "timestamp without time zone"] # 解析时间列 for col in time_cols: df[col] = pd.to_datetime(df[col]) # 转换剩余列的类型 non_time_map = {col: target_type_map[col] for col in target_type_map if col not in time_cols} df = df.astype(non_time_map)
注意事项
- 若某列当前类型不在
type_conversion中,会自动跳过该列的转换 - 转换前确保列数据无格式错误(比如布尔列不要混入非True/False值),否则会抛出转换异常
内容的提问来源于stack exchange,提问作者tlorel
相关产品推荐
相关产品推荐

