如何在Pandas DataFrame中筛选degCent降幅超70%后趋平的unitID
筛选满足温度降幅与趋平条件的unitID解决方案
数据预处理
首先确保时间列格式正确并按时序排列,这是时间序列分析的基础:
import pandas as pd # 将日期时间列转换为Pandas datetime类型 df['Date/Time'] = pd.to_datetime(df['Date/Time'], format='%d/%m/%Y %H:%M:%S.%f') # 按测试ID、设备ID、时间排序,保证数据时序逻辑正确 df = df.sort_values(by=['testID', 'unitID', 'Date/Time']).reset_index(drop=True)
定义条件判断函数
针对每个unitID的时间序列,编写函数验证两个核心条件:
- 峰值到谷值的降幅超过70%
- 谷值后出现至少5秒的趋平状态(维持在谷值±5%范围内)
def meets_requirements(group): # 1. 找到该设备的温度峰值及对应位置 peak_temp = group['degCent'].max() peak_position = group['degCent'].idxmax() # 仅考虑峰值之后的温度数据 post_peak_data = group.loc[peak_position:] # 2. 找到峰值后的最低温度(谷值)及位置 valley_temp = post_peak_data['degCent'].min() valley_position = post_peak_data['degCent'].idxmin() # 计算降幅比例,低于70%直接不满足 drop_percentage = (peak_temp - valley_temp) / peak_temp if drop_percentage <= 0.7: return False # 3. 验证谷值后的趋平状态 post_valley_data = group.loc[valley_position:] # 计算谷值的±5%区间 lower_limit = valley_temp * 0.95 upper_limit = valley_temp * 1.05 # 标记每个时间点是否在区间内 within_range = post_valley_data['degCent'].between(lower_limit, upper_limit) # 计算连续处于区间内的最大时长(1秒间隔,连续True的数量即秒数) consecutive_counts = within_range.astype(int).groupby((within_range != within_range.shift()).cumsum()).sum() max_consecutive = consecutive_counts.max() if not consecutive_counts.empty else 0 # 满足连续5秒及以上则返回True return max_consecutive >= 5
筛选目标测试中的合格设备
以testID=215299311为例,筛选满足条件的unitID:
# 聚焦目标测试数据 target_test_data = df[df['testID'] == 215299311] # 按unitID分组验证条件 validation_result = target_test_data.groupby('unitID').apply(meets_requirements) # 提取合格的unitID列表 qualified_units = validation_result[validation_result].index.tolist() print("符合条件的unitID:", qualified_units)
说明与调整点
- 若峰值定义不是全局最大值,而是下降趋势起始前的局部峰值,可通过计算温度差分(
group['degCent'].diff())识别下降阶段的起点,再取起点前的最大值作为峰值。 - 趋平的时间判断基于1秒间隔的连续数据,若实际数据存在缺失,需先补全时间序列(使用
resample方法)再进行判断。
内容的提问来源于stack exchange,提问作者James
相关产品推荐
相关产品推荐

