pandas混用object与string类型列时的运算报错问题咨询
pandas 1.4.2版本StringDtype类型兼容性问题
问题背景
pandas官方推荐使用dtype('string')类型替代传统的dtype('O')(object类型)存储字符串数据,但在1.4.2版本中使用该类型时会触发两类非预期异常。
问题1:不同字符串类型生成的透视表无法执行算术运算
复现代码
import pandas as pd df = pd.DataFrame() df['a']=[1.0,2.0,3.0] df['b']=list('abc') df['c']=list('def') df['cs'] = df['c'].astype('string') p1 = df.pivot_table(index='b', columns='c', values='a') p2 = df.pivot_table(index='b', columns='cs', values='a') p1+p2 # 此处触发报错
异常信息
ValueError: Location based indexing can only have [integer, integer slice (START point is INCLUDED, END point is EXCLUDED), listlike of integers, boolean array] types
根因定位
两个透视表p1、p2形状完全一致,但列索引数据类型不匹配:
p1的列索引类型为dtype('O')(object类型)p2的列索引类型为string[python]
如果代码中混用dtype('O')和dtype('string')类型的字符串列,很容易触发这类非预期错误。
问题2:HDF5读写后String类型索引自动回退为object类型
复现示例
>>> p2.columns.dtype string[python] >>> p2.to_hdf('data.h5','p2') >>> p2b = pd.read_hdf('data.h5','p2') >>> p2b.columns.dtype dtype('O')
写入HDF5文件再重新读取后,原本为string[python]类型的列索引会自动回退为dtype('O')类型。
当前临时规避方案
目前临时处理方式是尽量避免将字符串列转换为dtype('string')类型,实际使用中dtype('O')类型的兼容性表现更好,dtype('string')在1.4.2版本下似乎暂不适合日常生产使用。
咨询:是否有其他开发者遇到过同类问题?该场景下的官方推荐处理方案是什么?
内容的提问来源于stack exchange,提问作者Chris
相关产品推荐
相关产品推荐

