Pandas如何快速将前缀为'comps'的列统一转换为float类型
问题原因
- 第一种
df.convert_dtypes()是pandas自动推断所有列的最优存储类型,不会单独针对comps开头的列做float转换,不符合需求。 - 第二种写法有两处错误:一是
df.columns是Index对象,调用字符串方法startswith需要先通过.str访问器,直接写df.columns.startswith本身就会报错;二是就算拿到了布尔数组,也没有索引到对应列的数据,直接调用astype(float)必然执行失败。
解决方案
第一步:筛选所有comps开头的列
comps_cols = df.columns[df.columns.str.startswith('comps')]
第二步:批量转换类型
最通用的写法,直接修改原数据对应列:
df[comps_cols] = df[comps_cols].astype(float)
如果数据中存在无法转换为float的脏数据,可以加errors参数处理:
# errors='coerce'表示把无法转成float的内容统一替换为NaN df[comps_cols] = df[comps_cols].astype(float, errors='coerce')
如果不想修改原df,或者需要链式操作避免SettingWithCopyWarning,可以用assign方法:
df = df.assign(**{col: df[col].astype(float, errors='coerce') for col in comps_cols})
内容的提问来源于stack exchange,提问作者pouchewar
相关产品推荐
相关产品推荐

