You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

累积时间序列数据异常检测:现有算法失效问题求解决方案

累积值时间序列的异常检测问题

我正在用机器学习做异常检测,当前核心挑战是处理累积值时间序列数据集:每个数据点代表累积总和,只会递增(比如今日值100,明日值会是150,不会递减)。我的目标是识别数据集中的异常值并精准定位异常发生的时间点。

数据划分方式

考虑时间序列特性,我没有用随机划分,而是把前80%的数据用作训练集,剩下的20%作为测试集。但因为数据是累积型的,测试集的所有值都高于训练集,完全超出了训练数据的取值范围。

现有方法的问题

我原本计划用IsolationForest和**LOF(局部离群点检测)**两种算法,但结果完全不符合预期:测试集的所有样本都被模型判定为异常(预测结果为-1),原因就是测试值全部超出了训练数据的取值范围(比如训练集值范围1-100,测试集是101-120)。

关键限制

我无法直接把累积值转换为原始增量值:开发阶段虽然能实现转换,但生产环境中每次预测只能拿到单条记录,要还原原始值必须访问前一条记录,这在生产场景里做不到。

示例代码

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.ensemble import IsolationForest

# Generate example data (same as before)
np.random.seed(42)
num_entries = 100
timestamps = pd.date_range(start='2023-01-01', periods=num_entries, freq='D')
cumulative_values = np.sort(np.random.randint(0, 1000, num_entries))
cumulative_values_with_anomalies = cumulative_values.copy()
cumulative_values_with_anomalies[20] = 1500  # Introduce an anomaly (For training dataset)
cumulative_values_with_anomalies[88] = 2000  # Introduce another anomaly (For test dataset)

# Create DataFrame
data = {
    'timestamp': timestamps,
    'value': cumulative_values_with_anomalies
}
df = pd.DataFrame(data)

# Detect anomalies
anomaly_indices = np.where(cumulative_values_with_anomalies > cumulative_values)[0]


# Plot the data 
plt.figure(figsize=(10, 6))
plt.plot(df['timestamp'], df['value'], marker='.', label='Data')
plt.scatter(df['timestamp'][anomaly_indices], df['value'][anomaly_indices], color='red', label='Anomaly')
plt.title('Cumulative Values and Anomalies (Isolation Forest)')
plt.xlabel('Timestamp')
plt.ylabel('Cumulative Value')
plt.xticks(rotation=45)
plt.grid(True)
plt.legend()
plt.tight_layout()
plt.show()


# Manually set the cutoff index for training and test data
cutoff_index = int(len(df) * 0.8)  # Use 80% of the data for training
train_df = df.iloc[:cutoff_index]
test_df = df.iloc[cutoff_index:]


# Reshape data for Isolation Forest
X_train = train_df['value'].values.reshape(-1, 1)

# Train Isolation Forest
clf = IsolationForest(contamination=0.1, random_state=42)  # Adjust contamination based on your data
clf.fit(X_train)

# Predict anomalies on test data
X_test = test_df['value'].values.reshape(-1, 1)
test_df['anomaly'] = clf.predict(X_test)
anomalies = test_df[test_df['anomaly'] == -1]

# Plot the data
plt.figure(figsize=(10, 6))
plt.plot(test_df['timestamp'], test_df['value'], marker='.', label='Data')
plt.scatter(anomalies['timestamp'], anomalies['value'], color='red', label='Anomaly')
plt.title('Cumulative Values and Anomalies (Isolation Forest)')
plt.xlabel('Timestamp')
plt.ylabel('Cumulative Value')
plt.xticks(rotation=45)
plt.grid(True)
plt.legend()
plt.tight_layout()
plt.show()

代码说明

上述代码生成了包含时间戳和累积值的DataFrame,并手动注入了异常值。用IsolationForest训练后,测试集的所有数据都被标记为异常,和实际预期不符。

恳请各位提供可行的解决方案?


内容的提问来源于stack exchange,提问作者Bennison J

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 10:53:21