pandas DataFrame字符串列转float时如何跳过无法转换的行?
问题解决方法
- 方法1:无法转换的内容转为缺失值NaN(推荐,可获得纯数值类型列)
使用pandas内置的pd.to_numeric()函数,设置errors='coerce'参数即可自动将无法转换为数值的内容替换为NaN:
import pandas as pd df['A'] = pd.to_numeric(df['A'], errors='coerce')
如果后续需要转换为整数类型,可使用pandas支持空值的Int64类型:
df['A'] = df['A'].astype('Int64')
- 方法2:无法转换的内容保留原始字符串,不做修改
自定义转换函数配合apply实现,仅对可转换的内容做类型转换:
def safe_convert(val): try: # 可根据需要改为int(val) return float(val) except (ValueError, TypeError): # 转换失败直接返回原值 return val df['A'] = df['A'].apply(safe_convert)
注意:该方法处理后列的数据类型仍为object,因为列中同时存在数值和字符串两种类型
内容的提问来源于stack exchange,提问作者el-cheapo
相关产品推荐
相关产品推荐

