按ID分组,基于多条件根据上一行值确定下一行值
基于上一行结果生成desired_output列的解决方案
示例数据
import pandas as pd sample_data = { 'id': [1,1,1,1,1,2,2,2,2,2], 'date_rank': [1,2,3,4,5,1,2,3,4,5], 'candidates': [1,0,0,3,0,0,0,0,2,0], 'desired_output':['New_filled','New_open','Double_open','Double_filled','New_open','New_open','Double_open','Double_open','Double_filled','New_open'] } df = pd.DataFrame(sample_data, columns=['id', 'date_rank','candidates', 'desired_output'])
期望输出结果:
id date_rank candidates desired_output 0 1 1 1 New_filled 1 1 2 0 New_open 2 1 3 0 Double_open 3 1 4 3 Double_filled 4 1 5 0 New_open 5 2 1 0 New_open 6 2 2 0 Double_open 7 2 3 0 Double_open 8 2 4 2 Double_filled 9 2 5 0 New_open
规则说明
- 每组(按
id分组)的第一条记录前缀固定为New,后缀由candidates决定:candidates=0为open,否则为filled - 后续记录的前缀由上一条记录的后缀决定:
- 若上一条后缀为
filled,当前前缀为New - 若上一条后缀为
open,当前前缀为Double
- 若上一条后缀为
- 所有记录的后缀均由当前行
candidates值判断,规则同上
解决方案代码
def generate_desired_output(group): output = [] # 处理组内第一条记录 first_cand = group.iloc[0]['candidates'] first_suffix = 'filled' if first_cand > 0 else 'open' output.append(f'New_{first_suffix}') # 处理组内后续记录 for idx in range(1, len(group)): prev_suffix = output[idx-1].split('_')[1] curr_cand = group.iloc[idx]['candidates'] curr_suffix = 'filled' if curr_cand > 0 else 'open' prefix = 'New' if prev_suffix == 'filled' else 'Double' output.append(f'{prefix}_{curr_suffix}') group['generated_output'] = output return group # 按id分组处理并合并结果 df = df.groupby('id', group_keys=False).apply(generate_desired_output) # 验证结果 print(df[['desired_output', 'generated_output']])
运行代码后,generated_output列会和原desired_output完全匹配。
逻辑说明
- 按
id分组确保每组规则独立生效,避免组间干扰 - 先处理每组第一条记录,固定前缀并生成对应后缀
- 遍历后续行时,通过上一行结果的后缀确定当前前缀,结合当前行候选人数量生成后缀,拼接成最终结果
- 使用
groupby+apply的方式,既保证逻辑清晰,又能高效处理分组数据
内容的提问来源于stack exchange,提问作者avgjoe13
相关产品推荐
相关产品推荐

