如何在Pandas DataFrame中生成班级内学生组合并计算feature列?
问题描述
现有如下结构的Pandas DataFrame df:
Class_ID Student_ID theta 6 1 0.2 6 2 0.2 6 4 0.1 6 3 0.5 3 2 0.1 3 5 0.2 3 7 0.22 3 4 0.4 3 9 0.08
已通过以下代码生成同Class_ID内Student_ID的两两组合DataFrame df_new:
df_new = df.merge( df, how='inner', on=['Class_ID'] ) df_new['Student_ID_x'] = df_new['Student_ID_x'].astype(int) df_new['Student_ID_y'] = df_new['Student_ID_y'].astype(int) df_new = df_new[df_new['Student_ID_x'] < df_new['Student_ID_y']] df_new['Student_Combination'] = [f'{x}_{y}' for x, y in zip(df_new['Student_ID_x'], df_new['Student_ID_y'])]
现在需要为df_new添加名为feature的列,通过以下自定义函数计算:
def func(theta_x, theta_y, *theta): numerator = theta_x + theta_y denominator = 0 for t in theta: denominator += t*t return numerator / denominator
其中*theta为同一Class_ID内除当前组合两个学生外的所有theta值(例如Class 6的Student_Combination=1_2,feature=(0.2+0.2)/(0.1²+0.5²)=1.53846154),目标输出包含对应feature列,如何实现?
解决方案
可以通过分组预计算结合行级函数实现,既保证逻辑准确又提升效率,步骤如下:
- 预计算班级theta平方总和
先按Class_ID分组,计算每个班级所有theta值的平方和,避免后续重复遍历计算:
class_sq_sum = df.groupby('Class_ID')['theta'].agg(lambda x: sum(t**2 for t in x)).reset_index(name='total_sq_sum')
- 合并班级信息到df_new
将预计算的班级平方和数据与df_new关联,让每行组合都能获取对应班级的总平方和:
df_new = df_new.merge(class_sq_sum, on='Class_ID', how='left')
- 定义feature计算函数
利用总平方和减去当前组合两个theta的平方,直接得到分母(总平方和 = theta_x² + theta_y² + 其余theta平方和),无需遍历剩余theta值:
def calc_feature(row): theta_x = row['theta_x'] theta_y = row['theta_y'] numerator = theta_x + theta_y denominator = row['total_sq_sum'] - (theta_x**2 + theta_y**2) # 处理分母为0的边界情况(如班级仅2个学生) return numerator / denominator if denominator != 0 else None
- 生成feature列
对df_new每行应用计算函数,生成目标列:
df_new['feature'] = df_new.apply(calc_feature, axis=1)
- 清理临时列
按需删除用于计算的临时列total_sq_sum:
df_new = df_new.drop(columns=['total_sq_sum'])
验证示例
以Class 6的Student_Combination=1_2为例:
- 班级总平方和 = 0.2² + 0.2² + 0.1² + 0.5² = 0.34
- 分母 = 0.34 - 0.2² - 0.2² = 0.26
- feature = (0.2+0.2)/0.26 ≈ 1.53846154,与示例结果完全一致。
完整可运行代码
import pandas as pd # 构造原始DataFrame data = { 'Class_ID': [6,6,6,6,3,3,3,3,3], 'Student_ID': [1,2,4,3,2,5,7,4,9], 'theta': [0.2,0.2,0.1,0.5,0.1,0.2,0.22,0.4,0.08] } df = pd.DataFrame(data) # 生成两两组合df_new df_new = df.merge(df, how='inner', on=['Class_ID']) df_new['Student_ID_x'] = df_new['Student_ID_x'].astype(int) df_new['Student_ID_y'] = df_new['Student_ID_y'].astype(int) df_new = df_new[df_new['Student_ID_x'] < df_new['Student_ID_y']] df_new['Student_Combination'] = [f'{x}_{y}' for x, y in zip(df_new['Student_ID_x'], df_new['Student_ID_y'])] # 预计算班级theta平方总和 class_sq_sum = df.groupby('Class_ID')['theta'].agg(lambda x: sum(t**2 for t in x)).reset_index(name='total_sq_sum') # 合并并计算feature df_new = df_new.merge(class_sq_sum, on='Class_ID', how='left') def calc_feature(row): theta_x = row['theta_x'] theta_y = row['theta_y'] numerator = theta_x + theta_y denominator = row['total_sq_sum'] - (theta_x**2 + theta_y**2) return numerator / denominator if denominator != 0 else None df_new['feature'] = df_new.apply(calc_feature, axis=1) df_new = df_new.drop(columns=['total_sq_sum']) # 输出结果 print(df_new)
内容的提问来源于stack exchange,提问作者Ishigami
相关产品推荐
相关产品推荐

