修复for循环生成权重后为Pandas DataFrame赋值加权列的错误
问题:衰减权重计算函数的加权列赋值错误
问题详情
我写了一个Python函数,用来计算衰减权重并生成科目加权分数列,但赋值到DataFrame时出现错误:每次循环的当前行权重会覆盖所有行的计算结果,最终只有最后一行的加权列是正确的,其余行都用了最后一次迭代的权重。
原函数代码
def get_decay_weights(df, stat, col_list): df = df.reset_index() results_dict = [] for i,row in df.iterrows(): year_numbers = len(row['Year']) max_stat = max(row[stat]) if max_stat == 0: equal_weights = 1/year_numbers weights = {f's{i+1}': equal_weights for i in range(year_numbers)} else: decay = {f's{i+1}': [] for i in range(year_numbers)} percent_stat = {f's{i+1}': [] for i in range(year_numbers)} if year_numbers >= 1: decay[list(decay)[-1]] = 1 percent_stat[list(percent_stat)[0]] = (row[stat][0]/max_stat) if year_numbers >= 2: decay[list(decay)[-2]] = 0.63 percent_stat[list(percent_stat)[1]] = (row[stat][1]/max_stat) if year_numbers >= 3: decay[list(decay)[-3]] = 0.63**2 percent_stat[list(percent_stat)[2]]= (row[stat][2]/max_stat) if year_numbers >= 4: decay[list(decay)[-4]] = 0.63**3 percent_stat[list(percent_stat)[3]] = (row[stat][3]/max_stat) cumulative_scores = {k: decay[k]*percent_stat[k] for k in decay} weights = {k:v/sum(cumulative_scores.values(), 0.0) for k,v in cumulative_scores.items()} for col in col_list: combined = [x * y for x, y in zip(list(weights.values()), list(row[col]))] print("Combined:", combined) df[f'{col}_weighted'] = df.apply( lambda row: [x * y for x, y in zip(list((weights.values())), list(row[col]))],axis=1) print(df[f'{col}_weighted'] ) return df df = get_decay_weights(df, stat = 'Studying Time', col_list=['Math', 'Science'])
错误现象
打印的combined结果是正确的,但赋值到df[f'{col}_weighted']时,lambda函数没有绑定当前行的weights,而是复用了循环中最后一次的weights变量,导致所有行都用最后一行的权重计算。
示例输入数据
| Person | Year | Studying Time | Math | English | Science |
|---|---|---|---|---|---|
| Steve | [1,2,3,4,5] | [400,478,512,517,810] | [95,93,94,89,92] | [96,97,92,82,83] | [92,94,97,93,80] |
| George | [1,2] | [379,500] | [89,91] | [92,87] | [75,82] |
| Charles | [1] | [545] | [87] | [89] | [92] |
| Andy | [1,2,3] | [510,560,801] | [92,94,97] | [89,89,82] | [79,78,91] |
计算逻辑
- 为每次考试尝试分配衰减权重:最后一次尝试权重为1,前一次为0.63,依次类推(最多5次尝试)
- 计算每次尝试的学习时间占最大学习时间的比例
- 将衰减权重与该比例相乘得到累积值
- 累积值除以总累积值得到最终权重
- 用权重乘以对应科目的分数列表,生成加权分数列(此步骤出错)
解决方案
问题根源
循环中使用df.apply(lambda ...)时,lambda函数不会立即绑定当前的weights变量,而是在实际执行时才去引用weights,此时循环已经结束,所有lambda都指向最后一次循环的weights。
修复后的函数
直接在循环中通过索引为当前行的加权列赋值,避免使用lambda引用外部变量:
def get_decay_weights(df, stat, col_list): df = df.reset_index(drop=True) # 重置索引并丢弃原索引,避免索引混乱 for idx, row in df.iterrows(): year_numbers = len(row['Year']) max_stat = max(row[stat]) # 计算当前行的权重 if max_stat == 0: weights = [1/year_numbers] * year_numbers else: # 生成衰减权重:最后一次为1,往前依次乘0.63 decay_weights = [0.63 ** (year_numbers - 1 - i) for i in range(year_numbers)] # 计算学习时间占比 stat_ratios = [x / max_stat for x in row[stat]] # 计算累积值并归一化得到最终权重 cumulative = [d * r for d, r in zip(decay_weights, stat_ratios)] total_cumulative = sum(cumulative) weights = [c / total_cumulative for c in cumulative] # 为当前行的每个科目计算加权分数并赋值 for col in col_list: weighted_scores = [w * s for w, s in zip(weights, row[col])] df.loc[idx, f'{col}_weighted'] = weighted_scores return df # 调用示例 df = get_decay_weights(df, stat='Studying Time', col_list=['Math', 'Science'])
关键改动说明
- 简化权重计算逻辑:用列表推导式替代原有的多个if分支,更简洁易维护,同时确保衰减权重的顺序正确(对应考试尝试的顺序)
- 直接按索引赋值:使用
df.loc[idx, f'{col}_weighted']直接为当前行的加权列赋值,避免了lambda引用外部变量的问题 - 重置索引时丢弃原索引:避免原索引混乱导致的赋值错误
内容的提问来源于stack exchange,提问作者bangbanggggg
相关产品推荐
相关产品推荐

