如何基于Isolation Forest输出的-1到1过渡段过滤时序数据
问题
处理时序数据集时,使用scikit-learn的IsolationForest识别异常值,输出列Isolation_forest取值为1(正常值)和-1(异常值)。需要提取首个从-1过渡到1、再转回-1的连续1片段:即从第一次出现-1→1的行开始,到下一次出现1→-1的前一行结束,过滤掉其他所有行。
示例数据
原始DataFrame:
| 时间 | X | Y | Z | Isolation_forest |
|---|---|---|---|---|
| 0 | 12 | 45 | 67 | -1 |
| 1 | 34 | 23 | 89 | -1 |
| 2 | 56 | 78 | 12 | 1 |
| 3 | 23 | 45 | 67 | 1 |
| 4 | 89 | 12 | 34 | 1 |
| 5 | 45 | 67 | 89 | 1 |
| 6 | 78 | 23 | 12 | -1 |
| 7 | 12 | 89 | 45 | -1 |
| 8 | 34 | 67 | 78 | -1 |
| 9 | 56 | 12 | 23 | 1 |
| 10 | 67 | 34 | 56 | 1 |
期望结果:
| 时间 | X | Y | Z | Isolation_forest |
|---|---|---|---|---|
| 2 | 56 | 78 | 12 | 1 |
| 3 | 23 | 45 | 67 | 1 |
| 4 | 89 | 12 | 34 | 1 |
| 5 | 45 | 67 | 89 | 1 |
此前尝试的代码会保留所有Isolation_forest=1的行,不符合需求:
df_filtered = df.loc[df['Isolation_forest'] == 1]
解决方案
通过标记状态转换、定位目标片段区间实现,步骤如下:
- 计算相邻行的
Isolation_forest差值,标记-1→1(差值为2)和1→-1(差值为-2)的转换点 - 找到第一个
-1→1的起始索引,以及该起始索引之后第一个1→-1的结束索引 - 提取两个索引之间的行(不包含结束转换点的行)
完整代码
import pandas as pd # 构造示例数据(实际使用时替换为你的数据集) data = { '时间': [0,1,2,3,4,5,6,7,8,9,10], 'X': [12,34,56,23,89,45,78,12,34,56,67], 'Y': [45,23,78,45,12,67,23,89,67,12,34], 'Z': [67,89,12,67,34,89,12,45,78,23,56], 'Isolation_forest': [-1,-1,1,1,1,1,-1,-1,-1,1,1] } df = pd.DataFrame(data) # 1. 计算相邻行的差值,识别状态转换点 df['diff'] = df['Isolation_forest'].diff() # 2. 定位第一个-1→1的起始索引 start_idx = df[df['diff'] == 2].index[0] # 3. 定位起始点之后第一个1→-1的结束索引 end_idx = df[(df['diff'] == -2) & (df.index > start_idx)].index[0] # 4. 提取目标区间的行并删除辅助列 result = df.loc[start_idx : end_idx-1].drop(columns='diff') print(result)
代码说明
df['Isolation_forest'].diff():计算当前行与上一行的差值,-1→1时差值为2,1→-1时差值为-2df[df['diff'] == 2].index[0]:获取第一次从异常转正常的行索引df[(df['diff'] == -2) & (df.index > start_idx)].index[0]:获取第一次正常转异常的行索引(确保在起始索引之后)df.loc[start_idx : end_idx-1]:提取从起始行到结束转换点的前一行,最后删除辅助列diff
边界情况处理
如果数据中不存在完整的-1→1→-1片段(比如最后一直是1),可添加判断避免报错:
# 检查是否存在有效起始和结束转换点 has_start = not df[df['diff'] == 2].empty has_end = not df[(df['diff'] == -2) & (df.index > start_idx)].empty if has_start else False if has_start and has_end: result = df.loc[start_idx : end_idx-1].drop(columns='diff') else: # 根据需求处理,例如返回空DataFrame result = pd.DataFrame(columns=df.columns.drop('diff'))
内容的提问来源于stack exchange,提问作者Ershad Ahamed
相关产品推荐
相关产品推荐

