pandas from_dict使用orient='index'行为不一致,属于Bug吗?
pandas 1.5.0版本from_dict混合类型索引报错原因解析
背景复现
你提供的两段代码核心差异为内层字典的键数量:
- 可正常运行的data1场景代码:
import pandas as pd data1 = { 2010 : {'width': 0.1, 'height' : 0.6}, 2011 : {'width': 0.2, 'height' : 0.7}, 2012 : {'width': 0.3, 'height' : 0.8}, '2010-2012' : {'width': 0.2, 'height' : 0.7} } pd.DataFrame.from_dict(data1, orient = 'index')
- 触发TypeError的data2场景代码:
data2 = { 2010 : {'width': 0.1}, 2011 : {'width': 0.2}, 2012 : {'width': 0.3}, '2010-2012' : {'width': 0.2} } pd.DataFrame.from_dict(data2, orient = 'index')
报错栈信息为:
~/.local/lib/python3.8/site-packages/pandas/core/indexes/api.py in union_indexes(indexes, sort) 184 result = indexes[0] 185 if isinstance(result, list): --> 186 result = Index(sorted(result)) 187 return result 188 TypeError: '<' not supported between instances of 'str' and 'int'
根本原因
这个差异是pandas 1.5.0版本的内部实现逻辑导致的:
- 当使用
orient='index'调用from_dict时,如果内层字典存在多个不同键(如data1有width、height两个键),pandas会先遍历所有内层字典完成列合并,此过程中外层字典的键(行索引)直接按传入顺序构造,不会触发排序逻辑,因此混合int和str类型的索引不会出现比较报错。 - 如果所有内层字典的键完全一致且只有一个(如data2所有内层都只有width一个键),pandas会走优化处理路径,直接批量转换值为二维数组,此过程中构造行索引时会默认对索引值执行排序操作,int类型的年份和str类型的
2010-2012无法比较大小,因此触发TypeError。
其他规避方案
除了你使用的转置方法外,也可以选择统一外层键的类型避免排序报错:
# 把所有外层键转为字符串类型后再构造DataFrame data2_str_keys = {str(k): v for k, v in data2.items()} pd.DataFrame.from_dict(data2_str_keys, orient='index')
内容的提问来源于stack exchange,提问作者the.real.gruycho
相关产品推荐
相关产品推荐

