You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Pandas中提取时间序列各分组最后连续Alarm=1的子数据集

解决方法:提取每个分组最后一段连续alarm=1的记录

Got it, let's work through this problem together. Your current code is just flagging all rows where alarm=1, but we need to isolate the last consecutive stretch of 1s for each category, and only include records before your target date. Here's how to do it step by step:

Step 1: 预处理数据(时间戳转换与筛选)

首先确保timestamp列是datetime格式,然后筛选出目标日期之前的所有记录:

import pandas as pd

# 如果timestamp还不是datetime类型,先转换
df['timestamp'] = pd.to_datetime(df['timestamp'])

# 定义目标日期(替换成你实际需要的日期)
target_date = pd.to_datetime('2024-05-20')
# 只保留目标日期之前的数据
filtered_df = df[df['timestamp'] < target_date]

Step 2: 识别连续的alarm块

核心思路是给每一段连续相同的alarm值打上唯一标识,这样就能区分开不同的连续序列:

def extract_last_continuous_alarm(group):
    # 生成连续块的唯一ID:当alarm值发生变化时,ID递增
    group['block_id'] = (group['alarm'] != group['alarm'].shift()).cumsum()
    
    # 筛选出所有alarm=1的块
    alarm_1_blocks = group[group['alarm'] == 1]
    
    # 如果该分组没有alarm=1的记录,返回空DataFrame
    if alarm_1_blocks.empty:
        return pd.DataFrame()
    
    # 获取最后一个alarm=1块的ID
    last_block_id = alarm_1_blocks['block_id'].iloc[-1]
    
    # 返回这个最后一块的所有记录,同时删除临时的block_id列
    return group[group['block_id'] == last_block_id].drop('block_id', axis=1)

Step 3: 应用到每个分组并合并结果

把上面的函数应用到每个category分组,再整理最终结果:

# 对每个category分组执行函数,然后重置索引
final_result = filtered_df.groupby('category').apply(extract_last_continuous_alarm).reset_index(drop=True)

为什么你原来的代码没生效

你写的df.groupby('category')['alarm'].apply(lambda x: x==1)只是生成了每个分组内alarm=1的布尔标记,它无法区分分散的连续序列——所以会把所有零散的alarm=1行都选出来,而不是我们需要的最后一段连续序列。

额外注意事项

  • 如果某个category在目标日期前没有alarm=1的记录,结果中会返回空行(你可以根据需求修改函数跳过这类分组)
  • 确认target_date的格式和timestamp列匹配,避免筛选出错
  • 这个方法支持任意长度的连续序列,不管是1分钟还是数小时的连续记录

内容的提问来源于stack exchange,提问作者Nithya Ramesh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 19:07:45