如何基于applicationnumber为Pandas DataFrame列值添加递增后缀
Pandas实现分组内列值添加递增序号后缀方案
核心实现逻辑
- 利用
groupby('applicationnumber').cumcount()生成每个分组内从0开始的自增序号,加1后得到从1开始的步骤编号 - 将目标列的原值和
_step_+步骤编号的字符串拼接,完成后缀添加
完整可运行代码
单列添加后缀示例
import pandas as pd # 构造示例数据 data_example = {'applicationnumber': ['XYZ104183736AA', 'XYZ104183736AA', 'XDASDHGHG54G', 'XDASDHGHG54G','XDASDHGHG54G'], 'event_name': ['verification', 'verification', 'verification', 'verification','verification'], 'working_time_in_seconds': [1000,2000,30000,10000,1004]} df_example = pd.DataFrame(data_example) # 生成分组内步骤序号 df_example['step_seq'] = df_example.groupby('applicationnumber').cumcount() + 1 # 为event_name列添加后缀 df_example['event_name'] = df_example['event_name'] + '_step_' + df_example['step_seq'].astype(str) # 可删除临时生成的序号列 df_example.drop('step_seq', axis=1, inplace=True)
多列批量添加后缀示例
如果需要同时给多个列添加后缀,可遍历目标列实现:
add_suffix_cols = ['event_name', 'working_time_in_seconds'] df_example['step_seq'] = df_example.groupby('applicationnumber').cumcount() + 1 for col in add_suffix_cols: # 非字符串类型列需要先转成字符串再拼接 df_example[col] = df_example[col].astype(str) + '_step_' + df_example['step_seq'].astype(str) df_example.drop('step_seq', axis=1, inplace=True)
处理效果说明
以示例数据为例,处理后applicationnumber为XYZ104183736AA的两条记录,event_name会分别变为verification_step_1、verification_step_2;applicationnumber为XDASDHGHG54G的三条记录,event_name会分别变为verification_step_1、verification_step_2、verification_step_3,完全匹配需求。
内容的提问来源于stack exchange,提问作者sebikooo
相关产品推荐
相关产品推荐

