如何获取DataFrame中ID=3首个连续子组的最后索引
问题:获取ID=3第一个连续重复子组的最后索引
我有一个DataFrame,其中ID为'3'的记录存在多次连续重复的情况。第一次连续重复形成的子组索引为2、3、4(索引从0开始),需要记录该子组的最后一个索引值:4。
示例数据
data = {'ID': ['1', '2', '3', '3', '3', '4', '2', '5', '3', '3', '6', '3', '3']} df = pd.DataFrame(data)
ID=3的3个连续子组为:
Index ID
2 3
3 3
4 3Index ID
8 3
9 3Index ID
11 3
12 3
我需要获取第一个子组的最后索引4,但尝试的代码只能拿到所有ID=3记录的最后索引12:
尝试的代码
import pandas as pd import numpy as np max_indx = ([]) for indx, row in df.iterrows(): if row['ID'] ==str(3): max_indx = np.append(max_indx, indx) max_element = np.amax(max_indx) print('index: ', max_element)
解决方案
方法1:分组标记法
通过标记连续相同ID的分组,精准定位第一个目标组的最后索引:
import pandas as pd data = {'ID': ['1', '2', '3', '3', '3', '4', '2', '5', '3', '3', '6', '3', '3']} df = pd.DataFrame(data) # 生成连续分组的唯一标识 df['group_id'] = (df['ID'] != df['ID'].shift()).cumsum() # 筛选ID=3的第一个分组,获取其最大索引 target_group_id = df[df['ID'] == '3']['group_id'].iloc[0] last_index = df[(df['ID'] == '3') & (df['group_id'] == target_group_id)].index.max() print('第一个连续子组的最后索引:', last_index)
运行结果:第一个连续子组的最后索引: 4
方法2:遍历终止法
直接遍历数据,找到第一个连续3子组的结束位置后立即停止,效率更高:
import pandas as pd data = {'ID': ['1', '2', '3', '3', '3', '4', '2', '5', '3', '3', '6', '3', '3']} df = pd.DataFrame(data) found_first_group = False last_idx = None for idx in range(len(df)): if df.loc[idx, 'ID'] == '3': # 判断是否是当前连续组的最后一个元素 if (idx == len(df)-1) or (df.loc[idx+1, 'ID'] != '3'): if not found_first_group: last_idx = idx found_first_group = True break # 找到目标后直接终止遍历 print('第一个连续子组的最后索引:', last_idx)
运行结果同样输出4,符合需求。
内容的提问来源于stack exchange,提问作者Sara
相关产品推荐
相关产品推荐

