使用Python统计时间序列中的唯一Power循环
循环统计逻辑的Python实现
原始数据
现有时间序列数据框df1如下:
DateTime Power[HP] PowerJump PowerStart 01/01/2018 0 0 0 01/02/2018 10 0 1 01/03/2018 99 1 0 01/04/2018 20 0 0 01/05/2018 60 0 0 01/06/2018 70 0 0 01/07/2018 15 0 0 01/08/2018 95 1 0 01/09/2018 0 0 0 01/10/2018 10 0 1 01/11/2018 0 0 0 01/12/2018 95 1 1
字段说明:
PowerJump:标记功率至少提升20HP且达到90HP及以上的情况PowerStart:标记发动机开始产生功率的所有情况
需求说明
统计循环数量,循环定义规则:
- 满足
PowerStart=1或PowerJump=1时,视为一个循环的触发点 - 若
PowerStart=1的下一行是PowerJump=1,则二者算作同一个循环,仅在PowerStart=1的行标记Cycle=1,后续的PowerJump=1行标记Cycle=0
本样本预期输出如下:
DateTime Power[HP] PowerJump PowerStart Cycle 01/01/2018 0 0 0 0 01/02/2018 10 0 1 1 01/03/2018 99 1 0 0 01/04/2018 20 0 0 0 01/05/2018 60 0 0 0 01/06/2018 70 0 0 0 01/07/2018 15 0 0 0 01/08/2018 95 1 0 1 01/09/2018 0 0 0 0 01/10/2018 10 0 1 1 01/11/2018 0 0 0 0 01/12/2018 95 1 1 1
Python实现代码
import pandas as pd # 构造原始数据 data = { 'DateTime': ['01/01/2018', '01/02/2018', '01/03/2018', '01/04/2018', '01/05/2018', '01/06/2018', '01/07/2018', '01/08/2018', '01/09/2018', '01/10/2018', '01/11/2018', '01/12/2018'], 'Power[HP]': [0, 10, 99, 20, 60, 70, 15, 95, 0, 10, 0, 95], 'PowerJump': [0, 0, 1, 0, 0, 0, 0, 1, 0, 0, 0, 1], 'PowerStart': [0, 1, 0, 0, 0, 0, 0, 0, 0, 1, 0, 1] } df1 = pd.DataFrame(data) # 初始化Cycle列为0 df1['Cycle'] = 0 # 标记所有初始符合条件的行(PowerStart=1 或 PowerJump=1) mask = (df1['PowerStart'] == 1) | (df1['PowerJump'] == 1) df1.loc[mask, 'Cycle'] = 1 # 处理PowerStart后紧跟PowerJump的情况:将后续的PowerJump行的Cycle设为0 follow_jump_mask = (df1['PowerStart'] == 1) & (df1['PowerJump'].shift(-1) == 1) jump_indices = follow_jump_mask[follow_jump_mask].index + 1 df1.loc[jump_indices, 'Cycle'] = 0 print(df1.to_string(index=False))
代码说明
- 先构造原始数据并初始化
Cycle列为0 - 标记所有满足
PowerStart=1或PowerJump=1的行,将Cycle设为1 - 筛选出
PowerStart=1且下一行是PowerJump=1的行,找到它们的下一行索引,将这些行的Cycle重置为0,实现两个触发点合并为一个循环的逻辑 - 最后打印输出结果,与预期一致
内容的提问来源于stack exchange,提问作者Wizytor
相关产品推荐
相关产品推荐

