You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas:按分组生成序列并根据指定条件重置

问题描述

需要基于patient和schedule列生成result列,规则如下:

  • 条件1:patient列中患者变更时,序列需重置
  • 条件2:同一患者内,若schedule = '000',序列需重置

示例数据集:

df2 = pd.DataFrame({'patient': ['one', 'one', 'one', 'one','two', 'two','two','two','two'],    
                    'schedule': ['111', '111', '000', '111', '111', '000','111','111','111'],        
                    'date': ['11/20/2022', '11/22/2022', '11/23/2022', '11/8/2022', '11/9/2022', '11/14/2022','11/20/2022', '11/22/2022', '11/23/2022']})

期望生成的结果DataFrame(新增result列):

result = pd.DataFrame({'patient': ['one', 'one', 'one', 'one','two', 'two','two','two','two'],    
                    'schedule': ['111', '111', '000', '111', '111', '000','111','111','111'],        
                    'date': ['11/20/2022', '11/22/2022', '11/23/2022', '11/8/2022', '11/9/2022', '11/14/2022','11/20/2022', '11/22/2022', '11/23/2022'],
                    'result': ['1st_Time', '2nd_Time', 'Reset', '1st_Time', '1st_Time', 'Reset','1st_Time','2nd_Time','3rd_Time']})
解决方案

实现逻辑:

  1. 先划分分组:当患者切换或遇到schedule='000'时,生成新的分组ID,以此确定序列需要重置的区间
  2. 在每个分组内对非000的行进行连续计数
  3. 将计数值转换为对应的序数标签(1st、2nd、3rd等),000的行直接标记为Reset

代码实现

import pandas as pd

# 加载示例数据
df2 = pd.DataFrame({'patient': ['one', 'one', 'one', 'one','two', 'two','two','two','two'],    
                    'schedule': ['111', '111', '000', '111', '111', '000','111','111','111'],        
                    'date': ['11/20/2022', '11/22/2022', '11/23/2022', '11/8/2022', '11/9/2022', '11/14/2022','11/20/2022', '11/22/2022', '11/23/2022']})

# 生成分组ID:患者切换或schedule为000时触发分组更新
df2['group_id'] = (df2['patient'] != df2['patient'].shift()) | (df2['schedule'] == '000')
df2['group_id'] = df2['group_id'].cumsum()

# 每个分组内生成连续序列数
df2['seq'] = df2.groupby('group_id').cumcount() + 1

# 定义序数转换函数
def ordinal_to_label(n):
    suffix_map = {1: '1st', 2: '2nd', 3: '3rd'}
    return suffix_map.get(n, f"{n}th") + '_Time'

# 生成result列
df2['result'] = df2.apply(
    lambda row: 'Reset' if row['schedule'] == '000' else ordinal_to_label(row['seq']),
    axis=1
)

# 移除中间辅助列(可选)
df2 = df2.drop(['group_id', 'seq'], axis=1)

# 输出结果
print(df2)

运行上述代码后,得到的DataFrame与期望的result完全一致。

内容的提问来源于stack exchange,提问作者Murali

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 22:35:27