累积时间序列数据异常检测:现有算法失效问题求解决方案
累积值时间序列的异常检测问题
我正在用机器学习做异常检测,当前核心挑战是处理累积值时间序列数据集:每个数据点代表累积总和,只会递增(比如今日值100,明日值会是150,不会递减)。我的目标是识别数据集中的异常值并精准定位异常发生的时间点。
数据划分方式
考虑时间序列特性,我没有用随机划分,而是把前80%的数据用作训练集,剩下的20%作为测试集。但因为数据是累积型的,测试集的所有值都高于训练集,完全超出了训练数据的取值范围。
现有方法的问题
我原本计划用IsolationForest和**LOF(局部离群点检测)**两种算法,但结果完全不符合预期:测试集的所有样本都被模型判定为异常(预测结果为-1),原因就是测试值全部超出了训练数据的取值范围(比如训练集值范围1-100,测试集是101-120)。
关键限制
我无法直接把累积值转换为原始增量值:开发阶段虽然能实现转换,但生产环境中每次预测只能拿到单条记录,要还原原始值必须访问前一条记录,这在生产场景里做不到。
示例代码
import pandas as pd import numpy as np import matplotlib.pyplot as plt from sklearn.ensemble import IsolationForest # Generate example data (same as before) np.random.seed(42) num_entries = 100 timestamps = pd.date_range(start='2023-01-01', periods=num_entries, freq='D') cumulative_values = np.sort(np.random.randint(0, 1000, num_entries)) cumulative_values_with_anomalies = cumulative_values.copy() cumulative_values_with_anomalies[20] = 1500 # Introduce an anomaly (For training dataset) cumulative_values_with_anomalies[88] = 2000 # Introduce another anomaly (For test dataset) # Create DataFrame data = { 'timestamp': timestamps, 'value': cumulative_values_with_anomalies } df = pd.DataFrame(data) # Detect anomalies anomaly_indices = np.where(cumulative_values_with_anomalies > cumulative_values)[0] # Plot the data plt.figure(figsize=(10, 6)) plt.plot(df['timestamp'], df['value'], marker='.', label='Data') plt.scatter(df['timestamp'][anomaly_indices], df['value'][anomaly_indices], color='red', label='Anomaly') plt.title('Cumulative Values and Anomalies (Isolation Forest)') plt.xlabel('Timestamp') plt.ylabel('Cumulative Value') plt.xticks(rotation=45) plt.grid(True) plt.legend() plt.tight_layout() plt.show() # Manually set the cutoff index for training and test data cutoff_index = int(len(df) * 0.8) # Use 80% of the data for training train_df = df.iloc[:cutoff_index] test_df = df.iloc[cutoff_index:] # Reshape data for Isolation Forest X_train = train_df['value'].values.reshape(-1, 1) # Train Isolation Forest clf = IsolationForest(contamination=0.1, random_state=42) # Adjust contamination based on your data clf.fit(X_train) # Predict anomalies on test data X_test = test_df['value'].values.reshape(-1, 1) test_df['anomaly'] = clf.predict(X_test) anomalies = test_df[test_df['anomaly'] == -1] # Plot the data plt.figure(figsize=(10, 6)) plt.plot(test_df['timestamp'], test_df['value'], marker='.', label='Data') plt.scatter(anomalies['timestamp'], anomalies['value'], color='red', label='Anomaly') plt.title('Cumulative Values and Anomalies (Isolation Forest)') plt.xlabel('Timestamp') plt.ylabel('Cumulative Value') plt.xticks(rotation=45) plt.grid(True) plt.legend() plt.tight_layout() plt.show()
代码说明
上述代码生成了包含时间戳和累积值的DataFrame,并手动注入了异常值。用IsolationForest训练后,测试集的所有数据都被标记为异常,和实际预期不符。
恳请各位提供可行的解决方案?
内容的提问来源于stack exchange,提问作者Bennison J
相关产品推荐
相关产品推荐

