如何在Pandas按pgm_type分组生成指定规则的双计数器?
问题
需要在Pandas DataFrame中按pgm_type字段分组后生成两个计数器:
- counter1:从1到3循环取值;
- counter2:统计counter1重置为1的次数。
示例DataFrame df_sample
import pandas as pd df_sample = pd.DataFrame([[1,"pgm_1","type_A"], [1,"pgm_1","type_A"], [1,"pgm_1","type_A"], [2,"pgm_2","type_A"], [2,"pgm_2","type_A"], [2,"pgm_2","type_A"], [2,"pgm_2","type_A"], [2,"pgm_2","type_A"], [3,"pgm_3","type_B"], [3,"pgm_3","type_B"], [3,"pgm_3","type_B"], [3,"pgm_3","type_B"], [4,"pgm_4","type_B"] ]) df_sample.columns = ["pgm_id","pgm_name","pgm_type"] print(df_sample)
输出:
pgm_id pgm_name pgm_type 0 1 pgm_1 type_A 1 1 pgm_1 type_A 2 1 pgm_1 type_A 3 2 pgm_2 type_A 4 2 pgm_2 type_A 5 2 pgm_2 type_A 6 2 pgm_2 type_A 7 2 pgm_2 type_A 8 3 pgm_3 type_B 9 3 pgm_3 type_B 10 3 pgm_3 type_B 11 3 pgm_3 type_B 12 4 pgm_4 type_B
目标结果df_target
df_target = pd.DataFrame([[1,"pgm_1","type_A",1,1], [1,"pgm_1","type_A",2,1], [1,"pgm_1","type_A",3,1], [2,"pgm_2","type_A",1,2], [2,"pgm_2","type_A",2,2], [2,"pgm_2","type_A",3,2], [2,"pgm_2","type_A",1,3], [2,"pgm_2","type_A",2,3], [3,"pgm_3","type_B",1,1], [3,"pgm_3","type_B",2,1], [3,"pgm_3","type_B",3,1], [3,"pgm_3","type_B",1,2], [4,"pgm_4","type_B",1,2] ]) df_target.columns = ["pgm_id","pgm_name","pgm_type","counter1","counter2"] print(df_target)
输出:
pgm_id pgm_name pgm_type counter1 counter2 0 1 pgm_1 type_A 1 1 1 1 pgm_1 type_A 2 1 2 1 pgm_1 type_A 3 1 3 2 pgm_2 type_A 1 2 4 2 pgm_2 type_A 2 2 5 2 pgm_2 type_A 3 2 6 2 pgm_2 type_A 1 3 7 2 pgm_2 type_A 2 3 8 3 pgm_3 type_B 1 1 9 3 pgm_3 type_B 2 1 10 3 pgm_3 type_B 3 1 11 3 pgm_3 type_B 1 2 12 4 pgm_4 type_B 1 2
解决方案
通过分组计算行号、取模运算和累计求和即可实现需求,步骤如下:
- 按
pgm_type分组,为每组内的行分配从0开始的连续序号:
df_sample['row_num'] = df_sample.groupby('pgm_type').cumcount()
- 生成
counter1:利用行号对3取模后加1,实现1-3的循环取值:
df_sample['counter1'] = (df_sample['row_num'] % 3) + 1
- 生成
counter2:判断counter1是否为1,对每组内的该条件结果做累计求和,统计重置次数:
df_sample['counter2'] = df_sample.groupby('pgm_type')['counter1'].apply(lambda x: (x == 1).cumsum())
- 删除临时的
row_num列,得到最终结果:
df_sample = df_sample.drop('row_num', axis=1) print(df_sample)
执行后输出结果与目标df_target完全一致。
内容的提问来源于stack exchange,提问作者Darko37
相关产品推荐
相关产品推荐

