You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pandas from_dict使用orient='index'行为不一致,属于Bug吗?

pandas 1.5.0版本from_dict混合类型索引报错原因解析

背景复现

你提供的两段代码核心差异为内层字典的键数量:

  1. 可正常运行的data1场景代码:
import pandas as pd 

data1 = {
    2010 : {'width': 0.1, 'height' : 0.6},
    2011 : {'width': 0.2, 'height' : 0.7},
    2012 : {'width': 0.3, 'height' : 0.8},
    '2010-2012' : {'width': 0.2, 'height' : 0.7}
}

pd.DataFrame.from_dict(data1, orient = 'index')
  1. 触发TypeError的data2场景代码:
data2 = {
    2010 : {'width': 0.1},
    2011 : {'width': 0.2},
    2012 : {'width': 0.3},
    '2010-2012' : {'width': 0.2}
}

pd.DataFrame.from_dict(data2, orient = 'index')

报错栈信息为:

~/.local/lib/python3.8/site-packages/pandas/core/indexes/api.py in union_indexes(indexes, sort)
    184         result = indexes[0]
    185         if isinstance(result, list):
--> 186             result = Index(sorted(result))
    187         return result
    188 

TypeError: '<' not supported between instances of 'str' and 'int'

根本原因

这个差异是pandas 1.5.0版本的内部实现逻辑导致的:

  • 当使用orient='index'调用from_dict时,如果内层字典存在多个不同键(如data1有width、height两个键),pandas会先遍历所有内层字典完成列合并,此过程中外层字典的键(行索引)直接按传入顺序构造,不会触发排序逻辑,因此混合int和str类型的索引不会出现比较报错。
  • 如果所有内层字典的键完全一致且只有一个(如data2所有内层都只有width一个键),pandas会走优化处理路径,直接批量转换值为二维数组,此过程中构造行索引时会默认对索引值执行排序操作,int类型的年份和str类型的2010-2012无法比较大小,因此触发TypeError。

其他规避方案

除了你使用的转置方法外,也可以选择统一外层键的类型避免排序报错:

# 把所有外层键转为字符串类型后再构造DataFrame
data2_str_keys = {str(k): v for k, v in data2.items()}
pd.DataFrame.from_dict(data2_str_keys, orient='index')

内容的提问来源于stack exchange,提问作者the.real.gruycho

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 10:36:04