Pandas中apply导致int64自动转为float64的原因咨询
为何Pandas会自动将int64转换为float64?
我已查阅以下问题:
- 《Involuntary conversion of int64 to float64 in pandas》
- 《Unwanted automatic type conversion》
- 《Pandas Dtypes : float64 to 'Object' Conversion》
但据我理解,这些问题的场景都不如我的情况简单。我在Jupyter Lab中运行代码:
>>> df.dtypes cd_fndo int64 dif float64 dtype: object
可见数据类型为int64和float64,但应用恒等函数后类型发生了变化:
>>> df.apply(lambda x: x, axis=1).dtypes cd_fndo float64 dif float64 dtype: object
不过仅处理第一列时,int64类型保持不变:
>>> df.iloc[:, :1].apply(lambda x: x, axis=1).dtypes cd_fndo int64 dtype: object
能否有人解释这种类型变化的原因?
原因解释
当使用df.apply(lambda x: x, axis=1)逐行处理数据时,Pandas会将每一行转换为一个Series对象。由于单一行中同时存在int64和float64两种类型,Pandas需要为这个行Series选择一个能兼容所有元素的通用数据类型——float64可以无损容纳整数,而int64无法存储浮点数,因此会自动将整列的类型统一提升为float64,以保证每行数据类型的一致性。
而仅处理第一列时,每行只有int64类型的数据,不存在类型兼容问题,因此无需进行类型提升,类型保持int64不变。
内容的提问来源于stack exchange,提问作者Leonardo Maffei
相关产品推荐
相关产品推荐

