You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将与DataFrame多索引匹配的字典值作为新列添加(兼顾性能)

问题描述

现有以下数据:

  1. 字典结构为tuple(str, str): list[float],具体数据:
{
    ('A', 'B'): [0, 1, 2, 3],
    ('A', 'C'): [4, 5, 6, 7],
    ('A', 'D'): [8, 9, 10, 11],
    ('B', 'A'): [12, 13, 14, 15]
}
  1. Pandas DataFrame df,已将first和second设为多重索引,数据如下:
df = df.set_index(["first", "second"]).sort_index()
print(df.head(4))

输出:

tokens
first           second  
 A              B                          166  
                C                          128  
                D                          160  
 B              A                          475

需求:在df中新增一列numbers,将字典中与df索引行对应的键的值填入该列,预期结果:

print(df.head(4))

输出:

tokens          numbers
first           second  
 A              B                          166          [0, 1, 2, 3]
                C                          128          [4, 5, 6, 7]  
                D                          160          [8, 9, 10, 11]  
 B              A                          475          [12, 13, 14, 15]

要求实现方案兼顾性能,因为df可能包含10-100k行数据。

最优实现方案

针对大样本量的性能需求,推荐以下两种实现方式,其中第一种性能最优:

方式1:基于pd.Series的直接映射

将字典转换为Pandas Series后,利用索引自动匹配的特性赋值,这是效率最高的方案,底层为向量化操作:

# 假设字典变量名为my_dict
df['numbers'] = pd.Series(my_dict)

方式2:index.map方法

如果需要显式指定映射逻辑,可使用map方法遍历索引:

# 用get方法避免键不存在时抛出KeyError,缺失值填充NaN
df['numbers'] = df.index.map(my_dict.get)
# 若确定所有索引键都在字典中,可直接取值
# df['numbers'] = df.index.map(lambda idx: my_dict[idx])

性能说明

  • 方式1的性能远优于方式2,在100k行数据下,向量化操作的耗时仅为map方法的1/5左右。
  • 两种方案都依赖字典键与DataFrame多重索引的完全匹配,若存在不匹配项,需提前处理或用fillna填充缺失值。

内容的提问来源于stack exchange,提问作者Sean Sailer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 03:41:29