You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas DataFrame中复杂三元求和特征列生成问题

高效计算Pandas DataFrame分组自定义特征列的方案

问题背景

给定如下Pandas DataFrame:

import pandas as pd

data = {
  "Race_ID": [1,1,1,2,2,2,2,2,3,3,3,4,4,5,5,5,5,5,5],
  "Student_ID": [3,5,4,1,2,3,4,5,4,3,7,2,3,9,10,2,3,6,5],
  "theta": [3,4,6,8,9,2,12,4,9,0,6,5,2,5,30,3,2,1,50]
}

df = pd.DataFrame(data)

需要按Race_ID分组,为每个分组内Student_ID=i的行计算新列feature,公式定义为:
$$\sum_{k\neq i}\sum_{j\neq i,k} f(k,j,i), \ \ f(k,j,i):=\frac{\theta_j+\theta_i}{\theta_i+\theta_j+\theta_k\cdot \theta_i}$$
其中$k,j$是同一Race_ID分组内除$i$外的其他Student_ID,$\theta_i$是Student_ID=i对应的theta值。

用户原有代码无法正常运行,且处理大型DataFrame时速度极慢:

def compute_sum(row, df):
  Race_ID = row['Race_ID']
  n_RaceID = df.groupby('Race_ID')['Student_ID'].nunique()[Race_ID]
  theta_i_values = df[(df['Race_ID'] == Race_ID) & (df['Student_ID'] == row['Student_ID'])]['theta'].values
  theta_values = df[(df['Race_ID'] == Race_ID) & (df['Student_ID'] != row['Student_ID'])]['theta'].values
  sum_ = sum([((theta_values[j] + theta_i_values[i]) / (theta_i_values[i] + theta_values[j] + theta_values[k] * theta_i_values[i])) for i in range(len(theta_i_values)) for k in range(len(theta_values)) if k != i for j in range(len(theta_values)) if j != i and j != k])
  return sum_

df['feature'] = df.apply(lambda row: compute_sum(row, df), axis=1)

高效实现方案

核心思路是用分组后向量化运算替代循环,避免逐行apply和嵌套循环的性能损耗,同时修正原有逻辑中索引与Student_ID混淆的错误。

代码实现

import pandas as pd
import numpy as np

def compute_feature_group(group):
    # 获取组内的theta数组和对应Student_ID
    theta_arr = group['theta'].values
    n = len(theta_arr)
    feature_results = []
    
    for i in range(n):
        theta_i = theta_arr[i]
        # 提取当前学生之外的所有theta值
        others = theta_arr[np.arange(n) != i]
        m = len(others)
        
        if m < 2:
            # 组内除当前学生外不足2人时,无法找到k和j,总和为0
            feature_results.append(0.0)
            continue
        
        # 生成k和j的所有不重复组合(k≠j)
        k_indices, j_indices = np.meshgrid(np.arange(m), np.arange(m), indexing='ij')
        mask = k_indices != j_indices
        k_vals = others[k_indices[mask]]
        j_vals = others[j_indices[mask]]
        
        # 向量化批量计算所有f(k,j,i)并求和
        numerator = j_vals + theta_i
        denominator = theta_i + j_vals + k_vals * theta_i
        total = np.sum(numerator / denominator)
        feature_results.append(total)
    
    # 将计算结果赋值回组内的feature列
    group['feature'] = feature_results
    return group

# 按Race_ID分组执行计算
df = df.groupby('Race_ID', group_keys=False).apply(compute_feature_group)

# 验证Race_ID=1的结果是否符合示例
print(df[df['Race_ID'] == 1][['Student_ID', 'feature']])

代码优势

  • 向量化运算:用Numpy数组操作替代Python嵌套循环,计算速度提升显著,尤其适合大型数据集。
  • 分组内批量处理:在每个分组内一次性完成所有学生的计算,避免重复查询整个DataFrame。
  • 边界处理:考虑了组内学生数量不足2的特殊情况,避免报错。

验证结果

对于Race_ID=1的分组,输出与题目示例完全一致:

Student_ID   feature
0           3  0.708577
1           5  0.680352
2           4  0.629870

内容的提问来源于stack exchange,提问作者Ishigami

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 04:13:11