如何在DataFrame的apply()方法中传入完整DataFrame与当前行索引?
问题:在apply()中传入完整DataFrame与当前行索引,计算排除当前行后的相关性
需求:现有correlation_df如下,需要新增一列diff_correlations,每个行值为排除当前行后,score列与cosine列的皮尔逊相关系数,希望通过apply()方法实现,需传入完整DataFrame和当前行索引以排除该行。
| id | score | cosine |
|---|---|---|
| 1 | 100 | 0.8 |
| 2 | 75 | 0.7 |
| 3 | 50 | 0.4 |
| 4 | 25 | 0.05 |
原问题代码:
import numpy as np import pandas as pd score = np.array([100, 75, 50, 25]) cosine = np.array([.8, 0.7, 0.4, .05]) correlation_df = pd.DataFrame( { "score": score, "cosine": cosine, } ) # 计算整体相关性(非需求,仅示例) corr = correlation_df.corr().values[0, 1]
用户现有迂回解决方案(可优化):
def my_fuct(row): i = int(row["index"]) r = list(range(correlation_df.shape[0])) r.remove(i) subset = correlation_df.iloc[r, :].copy() subset = subset.set_index("index") return subset.corr().values[0, 1] # 修正原代码语法错误(多余的=) correlation_df["diff_correlations"] = correlation_df.apply(my_fuct, axis=1)
优化解决方案
核心思路
- 利用
apply(axis=1)处理每行时,row.name直接返回当前行的索引值,无需额外将索引转为列 - 通过
apply()的args参数传入完整DataFrame,避免依赖全局变量,降低代码耦合 - 使用
drop(current_idx)直接排除当前行,替代手动生成索引列表再删除的低效操作
优化后代码
import numpy as np import pandas as pd # 生成DataFrame score = np.array([100, 75, 50, 25]) cosine = np.array([.8, 0.7, 0.4, .05]) correlation_df = pd.DataFrame( { "score": score, "cosine": cosine, } ) # 定义计算函数:传入当前行、完整DataFrame def calc_corr_exclude_row(row, full_df): current_idx = row.name # 获取当前行索引 # 排除当前行后计算相关性 subset = full_df.drop(current_idx) return subset['score'].corr(subset['cosine']) # 调用apply,传入完整DataFrame作为参数 correlation_df['diff_correlations'] = correlation_df.apply( calc_corr_exclude_row, axis=1, args=(correlation_df,) # args需为元组,单个元素要加逗号 ) print(correlation_df)
代码说明
row.name:当apply(axis=1)时,每行的name属性就是该行的索引值,无需额外处理args=(correlation_df,):将完整DataFrame作为参数传递给自定义函数,避免函数依赖全局变量,代码更健壮full_df.drop(current_idx):直接删除指定索引的行,比手动生成索引列表效率更高,尤其当数据量较大时subset['score'].corr(subset['cosine']):直接计算两列的相关性,比corr().values[0,1]更直观易读
内容的提问来源于stack exchange,提问作者Atticus
相关产品推荐
相关产品推荐

