You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Pandas DataFrame的超指数增长排列求和实现方案

问题描述

给定如下Pandas DataFrame:

import pandas as pd

data = {
  "Race_ID": [2,2,2,2,2,5,5,5,5,5,5],
  "Student_ID": [1,2,3,4,5,9,10,2,3,6,5],
  "theta": [8,9,2,12,4,5,30,3,2,1,50]
}

df = pd.DataFrame(data)

需要新增feature列,规则如下:

  • 按Race_ID分组,对每组内的每个Student_ID(记为i),计算所有符合条件的函数f值之和作为该学生的feature。
  • 函数f的参数要求:k、j、l为同组内不等于i的Student_ID,且满足k≠i,j≠i、k,l≠k、j、i;theta_i为当前i对应的theta值。
  • 函数f定义:
def f(thetak, thetaj, thetai, *theta):
  prod = 1
  for t in theta:
    prod = prod * t
  return ((thetai + thetaj) / (thetai + thetaj + thetai * thetak)) * prod 

优化逻辑

原始求和项数量为(n-1)!(n为组内学生总数),计算效率极低。利用f的对称性优化:

  • 固定i、j、k时,其余参数的所有排列对应的f值完全相等,只需计算1项再乘以(n-3)!,求和项数量降至(n-1)(n-2),大幅降低计算量。

期望输出结果:

import pandas as pd

data = {
  "Race_ID": [2,2,2,2,2,5,5,5,5,5,5],
  "Student_ID": [1,2,3,4,5,9,10,2,3,6,5],
  "theta": [8,9,2,12,4,5,30,3,2,1,50],
  "feature": [299.1960138012742, 268.93506341257876, 634.7909309816431, 204.18901708653254, 483.7234700875771, 53588.197759, 9395.539167178009, 78005.26224935807, 92907.8753942894, 118315.38359654899, 5600.243276203378]
}

df = pd.DataFrame(data)
实现方案

以下代码基于优化逻辑实现,确保计算效率与结果准确性:

import pandas as pd
import math
from itertools import combinations

def f(thetak, thetaj, thetai, *theta):
    prod = 1
    for t in theta:
        prod *= t
    return ((thetai + thetaj) / (thetai + thetaj + thetai * thetak)) * prod

def calculate_feature(group, student_id):
    # 获取当前学生的theta值
    thetai = group.loc[group['Student_ID'] == student_id, 'theta'].iloc[0]
    # 提取组内其他学生的theta列表
    other_thetas = group.loc[group['Student_ID'] != student_id, 'theta'].tolist()
    n_others = len(other_thetas)
    
    # 当组内除当前学生外不足3人时,无符合条件的组合,返回0
    if n_others < 3:
        return 0.0
    
    # 计算剩余参数的排列数:(n-3)!
    permutation_factor = math.factorial(n_others - 3)
    total = 0.0
    
    # 遍历所有j和k的不重复组合(j≠k)
    for j_idx, k_idx in combinations(range(n_others), 2):
        thetaj = other_thetas[j_idx]
        thetak = other_thetas[k_idx]
        # 获取除j、k外的剩余theta值
        remaining_thetas = [
            other_thetas[m] for m in range(n_others) 
            if m != j_idx and m != k_idx
        ]
        # 计算单组f值并乘以排列因子,累加到总和
        total += f(thetak, thetaj, thetai, *remaining_thetas) * permutation_factor
    
    return total

# 按行计算feature列
df['feature'] = df.apply(
    lambda row: calculate_feature(df[df['Race_ID'] == row['Race_ID']], row['Student_ID']),
    axis=1
)

# 验证结果
print(df)

代码说明

  1. calculate_feature函数:针对单个学生完成feature计算:
    • 先提取当前学生的theta_i和组内其他学生的theta集合;
    • 计算(n-3)!作为排列因子,对应剩余参数的所有排列数;
    • 使用combinations生成j和k的所有不重复组合,避免重复计算排列情况,对每个组合计算剩余theta列表,代入f函数后乘以排列因子,累加得到最终总和。
  2. DataFrame应用:通过apply方法按行调用计算函数,完成全量数据的feature列生成。

内容的提问来源于stack exchange,提问作者Ishigami

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 19:14:51