Pandas:按分组生成序列并根据指定条件重置
问题描述
需要基于patient和schedule列生成result列,规则如下:
- 条件1:
patient列中患者变更时,序列需重置 - 条件2:同一患者内,若
schedule = '000',序列需重置
示例数据集:
df2 = pd.DataFrame({'patient': ['one', 'one', 'one', 'one','two', 'two','two','two','two'], 'schedule': ['111', '111', '000', '111', '111', '000','111','111','111'], 'date': ['11/20/2022', '11/22/2022', '11/23/2022', '11/8/2022', '11/9/2022', '11/14/2022','11/20/2022', '11/22/2022', '11/23/2022']})
期望生成的结果DataFrame(新增result列):
result = pd.DataFrame({'patient': ['one', 'one', 'one', 'one','two', 'two','two','two','two'], 'schedule': ['111', '111', '000', '111', '111', '000','111','111','111'], 'date': ['11/20/2022', '11/22/2022', '11/23/2022', '11/8/2022', '11/9/2022', '11/14/2022','11/20/2022', '11/22/2022', '11/23/2022'], 'result': ['1st_Time', '2nd_Time', 'Reset', '1st_Time', '1st_Time', 'Reset','1st_Time','2nd_Time','3rd_Time']})
解决方案
实现逻辑:
- 先划分分组:当患者切换或遇到
schedule='000'时,生成新的分组ID,以此确定序列需要重置的区间 - 在每个分组内对非
000的行进行连续计数 - 将计数值转换为对应的序数标签(1st、2nd、3rd等),
000的行直接标记为Reset
代码实现
import pandas as pd # 加载示例数据 df2 = pd.DataFrame({'patient': ['one', 'one', 'one', 'one','two', 'two','two','two','two'], 'schedule': ['111', '111', '000', '111', '111', '000','111','111','111'], 'date': ['11/20/2022', '11/22/2022', '11/23/2022', '11/8/2022', '11/9/2022', '11/14/2022','11/20/2022', '11/22/2022', '11/23/2022']}) # 生成分组ID:患者切换或schedule为000时触发分组更新 df2['group_id'] = (df2['patient'] != df2['patient'].shift()) | (df2['schedule'] == '000') df2['group_id'] = df2['group_id'].cumsum() # 每个分组内生成连续序列数 df2['seq'] = df2.groupby('group_id').cumcount() + 1 # 定义序数转换函数 def ordinal_to_label(n): suffix_map = {1: '1st', 2: '2nd', 3: '3rd'} return suffix_map.get(n, f"{n}th") + '_Time' # 生成result列 df2['result'] = df2.apply( lambda row: 'Reset' if row['schedule'] == '000' else ordinal_to_label(row['seq']), axis=1 ) # 移除中间辅助列(可选) df2 = df2.drop(['group_id', 'seq'], axis=1) # 输出结果 print(df2)
运行上述代码后,得到的DataFrame与期望的result完全一致。
内容的提问来源于stack exchange,提问作者Murali
相关产品推荐
相关产品推荐

