如何对DataFrame中符合Cycle阈值的分组ID重复最后一行至Cycle=4?
需求实现:按条件扩展DataFrame分组
原始DataFrame构建代码
# 导入所需库 import pandas as pd # 创建数据集 data = {'id': [1, 1, 1, 1, 1,1, 1, 1, 1, 1, 1, 2, 2, 3, 3, 3, 3, 3, 3, 4, 5, 5, 5, 5, 5,5, 5, 5,5], 'cycle': [1,2, 3, 4, 5,6,7,8,9,10,11, 1,2, 1,2, 3, 4, 5,6, 1, 1,2, 3, 4, 5,6,7,8,9,], 'Salary': [7, 7, 7,8,9,10,11,12,13,14,15, 4, 5, 8,9,10,11,12,13, 8, 7, 7,9,10,11,12,13,14,15,], 'Children': ['No', 'Yes', 'Yes', 'Yes', 'Yes', 'No','No', 'Yes', 'Yes', 'Yes', 'No', 'Yes', 'No', 'No','Yes', 'Yes', 'No','No', 'Yes', 'Yes', 'No', 'Yes', 'No', 'No', 'Yes', 'Yes', 'Yes', 'Yes', 'No',], 'Days': [123, 128, 66, 66, 120, 141, 52,96, 120, 141, 52, 96, 120, 15,123, 128, 66, 120, 141, 141, 123, 128, 66, 123, 128, 66, 120, 141, 52,], } # 转换为DataFrame df = pd.DataFrame(data) print("df = \n", df)
需求说明
每个id对应一组递增的cycle值,各id的最大cycle分别为:id-1=11、id-2=2、id-3=6、id-4=1、id-5=9。
设定阈值cycle_threshold = 3,需完成以下处理:
- 若某个
id的最大cycle≤ 阈值,则重复该id分组的最后一行,直到cycle值达到4; - 其余
id分组保持原数据不变。
实现代码
cycle_threshold = 3 target_cycle = 4 # 按id分组处理每个子数据集 processed_groups = [] for id_val, group in df.groupby('id'): max_cycle = group['cycle'].max() if max_cycle <= cycle_threshold: # 复制分组的最后一行作为模板 last_row = group.iloc[-1:].copy() # 生成需要补充的cycle值对应的行 for cycle_val in range(max_cycle + 1, target_cycle + 1): new_row = last_row.copy() new_row['cycle'] = cycle_val processed_groups.append(new_row) # 添加原分组数据 processed_groups.append(group) else: # 无需补充的分组直接加入结果 processed_groups.append(group) # 合并所有分组并重置索引,再按id和cycle排序保证顺序正确 result_df = pd.concat(processed_groups, ignore_index=True) result_df = result_df.sort_values(by=['id', 'cycle']).reset_index(drop=True) print("处理后的DataFrame:\n", result_df)
代码逻辑说明
- 遍历每个
id对应的分组,计算该分组的最大cycle值; - 若最大
cycle≤ 阈值,基于分组最后一行生成从max_cycle+1到4的新行,将新行和原分组都加入结果列表; - 无需补充的分组直接加入结果列表;
- 合并所有分组后按
id和cycle排序,保证数据顺序符合预期。
内容的提问来源于stack exchange,提问作者NN_Developer
相关产品推荐
相关产品推荐

