Pandas如何按id分组有条件删除DataFrame中特定序列的行
解决办法
原有代码中head()/tail()只能提取固定数量的行,无法识别连续匹配的范围,因此无法批量删除开头/末尾的连续匹配行。可以按id分组后逐组判断连续行的边界实现需求,以下是可行实现:
import pandas as pd df = pd.DataFrame(data={ 'id': [3, 3, 3, 3, 3, 5, 5, 5, 5, 6, 6, 6, 6, 6, 6, 6, 6, 6], 'outcome': ['no', 'no', 'no', 'yes', 'no', 'no', 'no', 'yes', 'yes', 'no', 'no', 'yes', 'yes', 'yes', 'yes', 'yes', 'no', 'no'] }) def group_filter(g): # 1. 过滤分组开头所有连续的no first_yes = (g['outcome'] == 'yes').idxmax() # 分组内无yes直接返回空 if g.loc[first_yes, 'outcome'] != 'yes': return g.iloc[0:0] g = g.loc[first_yes:] # 2. 过滤分组末尾所有连续的yes last_not_yes = (g['outcome'] != 'yes') if last_not_yes.any(): last_not_yes_idx = last_not_yes[::-1].idxmax() g = g.loc[:last_not_yes_idx] else: # 剩余全为yes直接返回空 return g.iloc[0:0] return g df = df.groupby('id', group_keys=False).apply(group_filter) print(df)
运行后输出结果完全符合预期:
| index | id | outcome |
|---|---|---|
| 3 | 3 | yes |
| 4 | 3 | no |
| 11 | 6 | yes |
| 12 | 6 | yes |
| 13 | 6 | yes |
| 14 | 6 | yes |
| 15 | 6 | yes |
| 16 | 6 | no |
| 17 | 6 | no |
如果偏好向量化操作避免apply,也可以用累积和实现:
# 标记第一个yes及之后的行 df['keep_start'] = df.groupby('id')['outcome'].transform(lambda x: (x == 'yes').cumsum() >= 1) df = df[df['keep_start']] # 标记最后一个非yes及之前的行 df['keep_end'] = df.groupby('id')['outcome'].transform(lambda x: (x != 'yes')[::-1].cumsum() >= 1) df = df[df['keep_end']].drop(columns=['keep_start', 'keep_end'])
内容的提问来源于stack exchange,提问作者Ze0ruso
相关产品推荐
相关产品推荐

