如何筛选DataFrame按id分组后val列连续出现3个及以上w的对应id?
实现方法(Python Pandas)
第一步:构造测试数据
import pandas as pd df = pd.DataFrame({ 'id': ['a','a','a','a','b','b','b','c','c','d','d','d','d'], 'val': ['w','w','l','w','w','w','w','w','l','w','w','w','w'] })
第二步:核心筛选逻辑
这里提供两种常用实现方案:
方案1:连续序列计数法(逻辑清晰,性能更优)
def check_consecutive(series, min_cnt=3): # 遇到非w时生成新的分组标记,重置计数 group_flag = series.ne('w').cumsum() # 统计每段连续w的长度 consecutive_len = series.eq('w').groupby(group_flag).cumsum() # 判断是否存在符合长度要求的连续段 return consecutive_len.max() >= min_cnt # 按id分组筛选,去重得到结果 result = df.groupby('id').filter(lambda x: check_consecutive(x['val']))['id'].drop_duplicates().to_frame()
方案2:滚动窗口法(代码更简洁)
# 开大小为3的滚动窗口,判断是否存在全为w的窗口 result = df.groupby('id').filter( lambda x: x['val'].rolling(3).apply(lambda y: all(y == 'w')).max() == 1 )['id'].drop_duplicates().to_frame()
输出结果
执行以上代码后打印result,即可得到预期输出:
id 0 b 1 d
内容的提问来源于stack exchange,提问作者skulldoger
相关产品推荐
相关产品推荐

