遍历Pandas DataFrame行,按条件计算新变量val
Pandas按规则生成新列
val的实现 问题背景
现有如下Pandas DataFrame(约1000条数据):
import pandas as pd dict1 = {'category': {0: 0.0, 1: 1.0, 2: 0.0, 3: 1.0, 4: 0.0}, 'Id': {0: 24108, 1: 24307, 2: 24307, 3: 24411, 4: 24411}, 'count': {0: 3, 1: 2, 2: 33, 3: 98, 4: 33}, 'weight': {0: 0.5, 1: 0.2, 2: 0.7, 3: 1.2, 4: 0.39}} df1 = pd.DataFrame(dict1)
数据中部分Id仅对应1个category标签(如Id 24108),部分Id对应2个category标签(如Id 24307、24411)。需要生成新列val,规则如下:
- 若
Id仅关联1个标签:val = count * weight - 若
Id关联2个标签:- 比较该
Id下两条记录的count值 - count较小的记录:
val = count * (weight + 1) - count较大的记录:
val = count * (1 - weight)
- 比较该
预期输出:
category Id count weight val 0 0.0 24108 3 0.50 1.50 1 1.0 24307 2 0.20 2.40 2 0.0 24307 33 0.70 9.90 3 1.0 24411 98 1.20 -19.60 4 0.0 24411 33 0.39 42.90
实现代码
通过分组结合自定义函数实现,步骤如下:
import pandas as pd # 定义处理单个Id分组的逻辑 def calculate_val(group): if len(group) == 1: # 单条记录直接计算乘积 group['val'] = group['count'] * group['weight'] return group else: # 找出当前Id下count最大和最小的行 max_idx = group['count'].idxmax() min_idx = group['count'].idxmin() # 分别计算对应val值 group.loc[max_idx, 'val'] = group.loc[max_idx, 'count'] * (1 - group.loc[max_idx, 'weight']) group.loc[min_idx, 'val'] = group.loc[min_idx, 'count'] * (group.loc[min_idx, 'weight'] + 1) return group # 按Id分组应用函数,保留原始索引顺序 result_df = df1.groupby('Id', group_keys=False).apply(calculate_val).loc[df1.index] # 格式化输出保留两位小数 print(result_df.round(2))
验证结果
运行代码后输出与预期完全一致:
category Id count weight val 0 0.0 24108 3 0.50 1.50 1 1.0 24307 2 0.20 2.40 2 0.0 24307 33 0.70 9.90 3 1.0 24411 98 1.20 -19.60 4 0.0 24411 33 0.39 42.90
内容的提问来源于stack exchange,提问作者A.N.
相关产品推荐
相关产品推荐

