You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame中生成班级内学生组合并计算feature列?

问题描述

现有如下结构的Pandas DataFrame df:

Class_ID    Student_ID  theta
6           1           0.2
6           2           0.2
6           4           0.1
6           3           0.5
3           2           0.1
3           5           0.2
3           7           0.22
3           4           0.4
3           9           0.08

已通过以下代码生成同Class_ID内Student_ID的两两组合DataFrame df_new:

df_new = df.merge(
    df,
    how='inner',
    on=['Class_ID']
)

df_new['Student_ID_x'] = df_new['Student_ID_x'].astype(int)
df_new['Student_ID_y'] = df_new['Student_ID_y'].astype(int)
df_new = df_new[df_new['Student_ID_x'] < df_new['Student_ID_y']]
df_new['Student_Combination'] = [f'{x}_{y}' for x, y in zip(df_new['Student_ID_x'], df_new['Student_ID_y'])]

现在需要为df_new添加名为feature的列,通过以下自定义函数计算:

def func(theta_x, theta_y, *theta):
 numerator = theta_x + theta_y
 denominator = 0
 for t in theta:
   denominator += t*t
 return numerator / denominator

其中*theta为同一Class_ID内除当前组合两个学生外的所有theta值(例如Class 6的Student_Combination=1_2,feature=(0.2+0.2)/(0.1²+0.5²)=1.53846154),目标输出包含对应feature列,如何实现?

解决方案

可以通过分组预计算结合行级函数实现,既保证逻辑准确又提升效率,步骤如下:

  1. 预计算班级theta平方总和
    先按Class_ID分组,计算每个班级所有theta值的平方和,避免后续重复遍历计算:
class_sq_sum = df.groupby('Class_ID')['theta'].agg(lambda x: sum(t**2 for t in x)).reset_index(name='total_sq_sum')
  1. 合并班级信息到df_new
    将预计算的班级平方和数据与df_new关联,让每行组合都能获取对应班级的总平方和:
df_new = df_new.merge(class_sq_sum, on='Class_ID', how='left')
  1. 定义feature计算函数
    利用总平方和减去当前组合两个theta的平方,直接得到分母(总平方和 = theta_x² + theta_y² + 其余theta平方和),无需遍历剩余theta值:
def calc_feature(row):
    theta_x = row['theta_x']
    theta_y = row['theta_y']
    numerator = theta_x + theta_y
    denominator = row['total_sq_sum'] - (theta_x**2 + theta_y**2)
    # 处理分母为0的边界情况(如班级仅2个学生)
    return numerator / denominator if denominator != 0 else None
  1. 生成feature列
    对df_new每行应用计算函数,生成目标列:
df_new['feature'] = df_new.apply(calc_feature, axis=1)
  1. 清理临时列
    按需删除用于计算的临时列total_sq_sum:
df_new = df_new.drop(columns=['total_sq_sum'])

验证示例

以Class 6的Student_Combination=1_2为例:

  • 班级总平方和 = 0.2² + 0.2² + 0.1² + 0.5² = 0.34
  • 分母 = 0.34 - 0.2² - 0.2² = 0.26
  • feature = (0.2+0.2)/0.26 ≈ 1.53846154,与示例结果完全一致。

完整可运行代码

import pandas as pd

# 构造原始DataFrame
data = {
    'Class_ID': [6,6,6,6,3,3,3,3,3],
    'Student_ID': [1,2,4,3,2,5,7,4,9],
    'theta': [0.2,0.2,0.1,0.5,0.1,0.2,0.22,0.4,0.08]
}
df = pd.DataFrame(data)

# 生成两两组合df_new
df_new = df.merge(df, how='inner', on=['Class_ID'])
df_new['Student_ID_x'] = df_new['Student_ID_x'].astype(int)
df_new['Student_ID_y'] = df_new['Student_ID_y'].astype(int)
df_new = df_new[df_new['Student_ID_x'] < df_new['Student_ID_y']]
df_new['Student_Combination'] = [f'{x}_{y}' for x, y in zip(df_new['Student_ID_x'], df_new['Student_ID_y'])]

# 预计算班级theta平方总和
class_sq_sum = df.groupby('Class_ID')['theta'].agg(lambda x: sum(t**2 for t in x)).reset_index(name='total_sq_sum')

# 合并并计算feature
df_new = df_new.merge(class_sq_sum, on='Class_ID', how='left')

def calc_feature(row):
    theta_x = row['theta_x']
    theta_y = row['theta_y']
    numerator = theta_x + theta_y
    denominator = row['total_sq_sum'] - (theta_x**2 + theta_y**2)
    return numerator / denominator if denominator != 0 else None

df_new['feature'] = df_new.apply(calc_feature, axis=1)
df_new = df_new.drop(columns=['total_sq_sum'])

# 输出结果
print(df_new)

内容的提问来源于stack exchange,提问作者Ishigami

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 17:05:19