访问pandas DataFrame行时单元格格式丢失问题求助
Pandas按行访问单元格时保留原始数据类型的解决方法
问题现象
当Pandas DataFrame仅包含数值类型列(整数+浮点数)时,通过loc/iloc取整行后,原本的整数类型会被自动转换为浮点数,丢失原始类型。示例如下:
import pandas dfA = pandas.DataFrame({ 'A': [111, 222, 333], 'B': [1.3, 2.4, 3.5], }) # 取首行后,A列的值变成111.0,类型为numpy.float64 print(dfA.loc[0]) print(type(dfA.loc[0].A))
但如果DataFrame包含字符串类型列,取整行时各列会保留原始类型:
dfB = pandas.DataFrame({ 'A': [111, 222, 333], 'B': [1.3, 2.4, 3.5], 'C': ['one', 'two', 'three'] }) # 此时A列的值为111,类型保持numpy.int64 print(dfB.loc[0]) print(type(dfB.loc[0].A))
原因分析
Pandas中取整行返回的是Series对象,而Series要求所有元素使用统一的数据类型:
- 当仅包含数值列时,整数会被向上转型为浮点数(浮点数可兼容整数,反之不行),整个Series的dtype变为
float64; - 当存在字符串列时,Series的dtype会变为
object,每个元素可独立保留原始类型。
解决方法
1. 直接访问单个单元格(推荐)
不需要取整行,直接定位目标单元格,直接返回原始类型:
print(dfA.loc[0, 'A']) # 输出111 print(type(dfA.loc[0, 'A'])) # 输出<class 'numpy.int64'>
2. 将行转换为字典
用to_dict()把整行转为字典,每个值保留原始类型:
row_dict = dfA.iloc[0].to_dict() print(row_dict['A']) # 输出111 print(type(row_dict['A'])) # 输出<class 'numpy.int64'>
3. 使用itertuples遍历行
通过itertuples()遍历行,返回的NamedTuple每个字段保留原始数据类型:
for row in dfA.itertuples(index=False): print(row.A, type(row.A)) # 输出111 <class 'int'>(转为Python原生int)
4. 手动指定Series的类型
如果必须取整行后处理,手动将Series各列转回原始类型:
row = dfA.loc[0].astype({'A': 'int64', 'B': 'float64'}) print(row.A) # 输出111 print(type(row.A)) # 输出<class 'numpy.int64'>
5. 创建DataFrame时指定object类型(不推荐)
创建时指定dtype=object,让列以object类型存储,取整行时不会转型,但会降低运算性能:
dfA = pandas.DataFrame({ 'A': [111, 222, 333], 'B': [1.3, 2.4, 3.5], }, dtype=object) print(dfA.loc[0].A) # 输出111 print(type(dfA.loc[0].A)) # 输出<class 'int'>
内容的提问来源于stack exchange,提问作者buhtz
相关产品推荐
相关产品推荐

