Pandas将object类型货币列转int时报float无replace属性错误如何处理
货币格式列转整数类型报错原因及解决方案
错误产生原因
- 尽管
Revenue列整体数据类型为object,但列中的空值NaN本质属于float类型,当apply遍历到NaN值时,对float类型调用字符串的replace方法,自然会触发AttributeError: 'float' object has no attribute 'replace'报错。 - 额外隐藏问题:原数据保留两位小数(如
$0.00),直接处理后转int64会因小数点存在报错;同时numpy原生的np.int64类型不支持空值,未处理空值直接转换也会失败。
正确实现方案
推荐使用pandas内置向量化操作实现,相比逐行apply执行效率更高,且自动兼容空值:
方案1:保留空值,使用pandas支持空值的Int64类型
如果后续计算需要保留空值标识,使用pandas专属的Int64(首字母大写)整数类型,该类型支持缺失值:
import pandas as pd import numpy as np if df['Revenue'].dtype == 'object': # 批量移除$符号和千分位逗号,str方法自动跳过NaN值 df['Revenue'] = df['Revenue'].str.replace(r'[$,]', '', regex=True) # 先转为浮点型,再转支持空值的Int64整数类型 df['Revenue'] = pd.to_numeric(df['Revenue'], errors='coerce').astype('Int64')
方案2:空值填充为0,使用numpy标准int64类型
如果空值在业务中可按0处理,直接填充后转标准int64即可:
if df['Revenue'].dtype == 'object': df['Revenue'] = df['Revenue'].str.replace(r'[$,]', '', regex=True) df['Revenue'] = pd.to_numeric(df['Revenue'], errors='coerce').fillna(0).astype(np.int64)
原有apply写法的修正方案(不推荐,仅作参考)
如果一定要沿用apply写法,只需在lambda中增加字符串类型判断,跳过空值即可:
if df['Revenue'].dtype == 'object': df['Revenue'] = df['Revenue'].apply(lambda x: x.replace('$','').replace(',','') if isinstance(x, str) else x) df['Revenue'] = pd.to_numeric(df['Revenue'], errors='coerce').astype('Int64')
内容的提问来源于stack exchange,提问作者kms
相关产品推荐
相关产品推荐

