Pandas中Lambda调用带多参数自定义函数报错问题求助
问题分析与解决方法
错误原因
你的WeightedScore函数仅定义了两个参数x和probs,但调用时传入了3个位置参数(x['A'], x['B'], x['C'])+ 关键字参数probs,导致probs被同时通过位置(第二个位置参数x['B'])和关键字传递,触发了参数重复赋值的错误。
解决方法
方法1:修正函数调用方式
将A/B/C列打包为一个序列传给x参数,让函数能通过索引获取对应值:
def WeightedScore(x, probs): total = 0 for i in range(len(probs)): total += x[i]*probs[i] return total df['out'] = df.apply(lambda x: WeightedScore(x[['A', 'B', 'C']], probs=(0.2, 0.3, 0.5)), axis=1)
方法2:优化函数,用numpy点积简化计算
替换手动循环,利用numpy的点积函数直接计算加权和:
import numpy as np def WeightedScore(x, probs): return np.dot(x, probs) df['out'] = df.apply(lambda x: WeightedScore(x[['A', 'B', 'C']], (0.2, 0.3, 0.5)), axis=1)
方法3:用向量运算替代apply(推荐,效率更高)
apply是逐行处理,大数据集下效率低,直接用pandas的向量运算实现:
weights = (0.2, 0.3, 0.5) # 写法1:直接列与权重相乘求和 df['out'] = df['A']*weights[0] + df['B']*weights[1] + df['C']*weights[2] # 写法2:更灵活的矩阵点积(适合多列场景) df['out'] = df[['A','B','C']].dot(weights)
内容的提问来源于stack exchange,提问作者DaCard
相关产品推荐
相关产品推荐

