You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何无循环实现Pandas DataFrame间的定向数据填充?

问题

我有两个Pandas DataFrame,想用其中一个的数据填充另一个。二者共用一套索引,但索引排列方式不同,而且第二个DataFrame不是简单的透视表,它的索引是为后续分析特意选定的。我已经用for循环实现了一个简单方案,但效率可能不高,想知道有没有不用循环的实现方式?

数据结构可视化

a    b    c                             values
a .    .    .                    a    d    .
b .    .    .                         g    .
c .    .    .                    b    h    .
d .    .    .                    c    e    .
e .    .    .                         d    .
f .    .    .            --->          j    .
g .    .    .
h .    .    .
i .    .    .
j .    .    .

现有循环实现代码

import numpy as np
import pandas as pd

np.random.seed(0) # 保证结果可复现
test = pd.DataFrame(np.random.rand(10,3), index=["a","b","c","d","e","f","g","h","i","j"], columns=["a","b","c"])
new_array = pd.DataFrame({"t1":["a","a","b","c","c","c"],"t2":["d","g","h","e","d","j"], "value":[np.nan, np.nan,np.nan,np.nan,np.nan,np.nan]})
new_array = new_array.set_index(["t1", "t2"])
for indexes in new_array.index:
    new_array.loc[indexes, "value"] = test.loc[indexes[1], indexes[0]]
无循环的高效实现方案

可以用Pandas内置方法实现批量匹配赋值,完全避免循环,效率显著提升。

方法一:使用lookup批量提取

lookup支持根据行、列标签对批量提取对应值,完美匹配需求:

# 从new_array的多层索引中拆分出需要的行、列标签
row_labels = new_array.index.get_level_values('t2')
col_labels = new_array.index.get_level_values('t1')

# 批量提取test中对应位置的值并赋值
new_array['value'] = test.lookup(row_labels, col_labels)

方法二:通过长格式转换+索引对齐

把test转成长格式后,直接通过索引对齐完成赋值:

# 将test转成长格式,构建与new_array匹配的索引
test_long = test.stack().reset_index()
test_long.columns = ['t2', 't1', 'value']
test_long = test_long.set_index(['t1', 't2'])

# 利用Pandas索引对齐特性直接赋值
new_array['value'] = test_long['value']

效果说明

两种方法都能得到和循环实现完全一致的结果,在处理大数据量时,效率比循环高数十倍甚至上百倍,数据规模越大,优势越明显。


内容的提问来源于stack exchange,提问作者linkey apiacess

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 13:13:21