You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过字典转换Python DataFrame中的数据类型?

用字典批量转换DataFrame数据类型的实现方法

核心思路

先基于你提供的类型映射字典,生成DataFrame各列对应的目标类型映射关系,再通过astype()方法批量转换;针对时间类型可单独处理,避免字符串转datetime64时出现解析失败问题。

实现步骤及代码

  1. 准备类型映射字典(你提供的)
type_conversion = {
    "bigint": "int64",
    "boolean": "bool",
    "character varying": "str",
    "double precision": "float64",
    "integer": "int32",
    "numeric": "float64",
    "timestamp without time zone": "datetime64"
}
  1. 生成列与目标类型的映射
    首先要明确DataFrame每列的当前数据类型名称(需与type_conversion的键匹配,比如从数据库读取的原始类型)。假设你已获取列与当前类型的对应关系(如col_types字典,键为列名,值为当前类型),生成目标映射:
# 仅保留可转换列的目标类型映射
target_type_map = {
    col: type_conversion[ctype] 
    for col, ctype in col_types.items() 
    if ctype in type_conversion
}

如果当前类型是从DataFrame的dtypes获取,可先转换为字符串匹配:

col_types = {col: str(df[col].dtype) for col in df.columns}
target_type_map = {
    col: type_conversion[ctype] 
    for col, ctype in col_types.items() 
    if ctype in type_conversion
}
  1. 批量转换数据类型
  • 普通类型直接用astype()转换:
df = df.astype(target_type_map)
  • 时间类型单独处理(适配字符串格式的时间数据):
    如果原始时间列是字符串格式,先通过pd.to_datetime()解析,再设置类型:
# 筛选出时间类型列
time_cols = [col for col, ctype in col_types.items() if ctype == "timestamp without time zone"]
# 解析时间列
for col in time_cols:
    df[col] = pd.to_datetime(df[col])
# 转换剩余列的类型
non_time_map = {col: target_type_map[col] for col in target_type_map if col not in time_cols}
df = df.astype(non_time_map)

注意事项

  • 若某列当前类型不在type_conversion中,会自动跳过该列的转换
  • 转换前确保列数据无格式错误(比如布尔列不要混入非True/False值),否则会抛出转换异常

内容的提问来源于stack exchange,提问作者tlorel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 16:03:27