基于Pandas DataFrame的超指数增长排列求和实现方案
问题描述
给定如下Pandas DataFrame:
import pandas as pd data = { "Race_ID": [2,2,2,2,2,5,5,5,5,5,5], "Student_ID": [1,2,3,4,5,9,10,2,3,6,5], "theta": [8,9,2,12,4,5,30,3,2,1,50] } df = pd.DataFrame(data)
需要新增feature列,规则如下:
- 按
Race_ID分组,对每组内的每个Student_ID(记为i),计算所有符合条件的函数f值之和作为该学生的feature。 - 函数
f的参数要求:k、j、l为同组内不等于i的Student_ID,且满足k≠i,j≠i、k,l≠k、j、i;theta_i为当前i对应的theta值。 - 函数
f定义:
def f(thetak, thetaj, thetai, *theta): prod = 1 for t in theta: prod = prod * t return ((thetai + thetaj) / (thetai + thetaj + thetai * thetak)) * prod
优化逻辑
原始求和项数量为(n-1)!(n为组内学生总数),计算效率极低。利用f的对称性优化:
- 固定i、j、k时,其余参数的所有排列对应的
f值完全相等,只需计算1项再乘以(n-3)!,求和项数量降至(n-1)(n-2),大幅降低计算量。
期望输出结果:
import pandas as pd data = { "Race_ID": [2,2,2,2,2,5,5,5,5,5,5], "Student_ID": [1,2,3,4,5,9,10,2,3,6,5], "theta": [8,9,2,12,4,5,30,3,2,1,50], "feature": [299.1960138012742, 268.93506341257876, 634.7909309816431, 204.18901708653254, 483.7234700875771, 53588.197759, 9395.539167178009, 78005.26224935807, 92907.8753942894, 118315.38359654899, 5600.243276203378] } df = pd.DataFrame(data)
实现方案
以下代码基于优化逻辑实现,确保计算效率与结果准确性:
import pandas as pd import math from itertools import combinations def f(thetak, thetaj, thetai, *theta): prod = 1 for t in theta: prod *= t return ((thetai + thetaj) / (thetai + thetaj + thetai * thetak)) * prod def calculate_feature(group, student_id): # 获取当前学生的theta值 thetai = group.loc[group['Student_ID'] == student_id, 'theta'].iloc[0] # 提取组内其他学生的theta列表 other_thetas = group.loc[group['Student_ID'] != student_id, 'theta'].tolist() n_others = len(other_thetas) # 当组内除当前学生外不足3人时,无符合条件的组合,返回0 if n_others < 3: return 0.0 # 计算剩余参数的排列数:(n-3)! permutation_factor = math.factorial(n_others - 3) total = 0.0 # 遍历所有j和k的不重复组合(j≠k) for j_idx, k_idx in combinations(range(n_others), 2): thetaj = other_thetas[j_idx] thetak = other_thetas[k_idx] # 获取除j、k外的剩余theta值 remaining_thetas = [ other_thetas[m] for m in range(n_others) if m != j_idx and m != k_idx ] # 计算单组f值并乘以排列因子,累加到总和 total += f(thetak, thetaj, thetai, *remaining_thetas) * permutation_factor return total # 按行计算feature列 df['feature'] = df.apply( lambda row: calculate_feature(df[df['Race_ID'] == row['Race_ID']], row['Student_ID']), axis=1 ) # 验证结果 print(df)
代码说明
calculate_feature函数:针对单个学生完成feature计算:- 先提取当前学生的
theta_i和组内其他学生的theta集合; - 计算
(n-3)!作为排列因子,对应剩余参数的所有排列数; - 使用
combinations生成j和k的所有不重复组合,避免重复计算排列情况,对每个组合计算剩余theta列表,代入f函数后乘以排列因子,累加得到最终总和。
- 先提取当前学生的
- DataFrame应用:通过
apply方法按行调用计算函数,完成全量数据的feature列生成。
内容的提问来源于stack exchange,提问作者Ishigami
相关产品推荐
相关产品推荐

