在Pandas中提取时间序列各分组最后连续Alarm=1的子数据集
解决方法:提取每个分组最后一段连续alarm=1的记录
Got it, let's work through this problem together. Your current code is just flagging all rows where alarm=1, but we need to isolate the last consecutive stretch of 1s for each category, and only include records before your target date. Here's how to do it step by step:
Step 1: 预处理数据(时间戳转换与筛选)
首先确保timestamp列是datetime格式,然后筛选出目标日期之前的所有记录:
import pandas as pd # 如果timestamp还不是datetime类型,先转换 df['timestamp'] = pd.to_datetime(df['timestamp']) # 定义目标日期(替换成你实际需要的日期) target_date = pd.to_datetime('2024-05-20') # 只保留目标日期之前的数据 filtered_df = df[df['timestamp'] < target_date]
Step 2: 识别连续的alarm块
核心思路是给每一段连续相同的alarm值打上唯一标识,这样就能区分开不同的连续序列:
def extract_last_continuous_alarm(group): # 生成连续块的唯一ID:当alarm值发生变化时,ID递增 group['block_id'] = (group['alarm'] != group['alarm'].shift()).cumsum() # 筛选出所有alarm=1的块 alarm_1_blocks = group[group['alarm'] == 1] # 如果该分组没有alarm=1的记录,返回空DataFrame if alarm_1_blocks.empty: return pd.DataFrame() # 获取最后一个alarm=1块的ID last_block_id = alarm_1_blocks['block_id'].iloc[-1] # 返回这个最后一块的所有记录,同时删除临时的block_id列 return group[group['block_id'] == last_block_id].drop('block_id', axis=1)
Step 3: 应用到每个分组并合并结果
把上面的函数应用到每个category分组,再整理最终结果:
# 对每个category分组执行函数,然后重置索引 final_result = filtered_df.groupby('category').apply(extract_last_continuous_alarm).reset_index(drop=True)
为什么你原来的代码没生效
你写的df.groupby('category')['alarm'].apply(lambda x: x==1)只是生成了每个分组内alarm=1的布尔标记,它无法区分分散的连续序列——所以会把所有零散的alarm=1行都选出来,而不是我们需要的最后一段连续序列。
额外注意事项
- 如果某个category在目标日期前没有
alarm=1的记录,结果中会返回空行(你可以根据需求修改函数跳过这类分组) - 确认
target_date的格式和timestamp列匹配,避免筛选出错 - 这个方法支持任意长度的连续序列,不管是1分钟还是数小时的连续记录
内容的提问来源于stack exchange,提问作者Nithya Ramesh
相关产品推荐
相关产品推荐

