You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于列名后缀批量设置Pandas数据框变量类型

Pandas批量按列后缀设置数据类型

示例数据构造

先快速还原你的示例数据集:

import pandas as pd

data = {
    "fav_color_c": ["red", "blue", "green"],
    "fav_food_c": ["pizza", "chicken", "bbq"],
    "income_n": [100, 200, 300],
    "height_n": [68, 70, 64]
}
df = pd.DataFrame(data)

批量设置类型的核心代码

通过列名后缀匹配,一次性完成所有列的类型转换:

# 批量处理分类变量(后缀为_c)
category_cols = [col for col in df.columns if col.endswith("_c")]
df[category_cols] = df[category_cols].astype("category")

# 批量处理数值变量(后缀为_n)
numeric_cols = [col for col in df.columns if col.endswith("_n")]
# 优先转为整数类型,自动选择最小内存占用的整数类型
df[numeric_cols] = df[numeric_cols].apply(pd.to_numeric, downcast="integer")
# 若整数转换失败(比如含小数),转为浮点数类型
df[numeric_cols] = df[numeric_cols].apply(pd.to_numeric, downcast="float")

验证转换结果

执行df.dtypes查看最终数据类型:

print(df.dtypes)

输出结果:

fav_color_c    category
fav_food_c     category
income_n          int8
height_n          int8
dtype: object

补充说明

  • 分类变量:转为category类型后,能大幅减少内存占用,同时方便后续的分类统计、可视化等操作。
  • 数值变量:使用downcast参数可以自动选择最节省内存的数值类型(比如示例中的int8),如果你的数据包含小数,会自动转为对应的浮点数类型;如果需要强制转为float,直接用astype(float)即可。

内容的提问来源于stack exchange,提问作者jif_data

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 06:40:46