You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

时序数据异常值标记失效:逻辑问题还是代码问题?

时序数据异常值检测问题:为何添加的异常未被识别?

尝试用Pandas简单检测时序数据异常值,在2024-04-01至2024-05-30时间段添加了大幅噪声,但代码未检测到这些异常值,请问是逻辑缺陷还是代码错误?

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt

baseline_price = 100.00
dates = pd.date_range(start='2024-01-01', end='2024-12-31', freq='D')
baseline_prices = pd.Series(baseline_price, index=dates)
baseline_prices += np.random.uniform(10, 15, size=len(baseline_prices))

shifted_prices = baseline_prices.copy()
shifted_prices.loc['2024-04-01':'2024-05-30'] += np.random.uniform(1000, 1015, size=60)

df = pd.DataFrame({'Date': dates, 'Stock Price': shifted_prices})

window_size, multiplier = 30, 2
rolling_mean = df['Stock Price'].rolling(window=window_size, center=True).mean()
rolling_std = df['Stock Price'].rolling(window=window_size, center=True).std()

upper_bound = rolling_mean + (multiplier * rolling_std)
lower_bound = rolling_mean - (multiplier * rolling_std)
outliers = df[(df['Stock Price'] > upper_bound) | (df['Stock Price'] < lower_bound)]

plt.figure(figsize=(10, 6))
plt.plot(df['Date'], df['Stock Price'], label='Stock Price')
plt.plot(df['Date'], rolling_mean, label=f'{window_size}-day Rolling Mean')
plt.fill_between(df['Date'], upper_bound, lower_bound, color='gray', alpha=0.2, label=f'{multiplier}-sigma Bounds')
plt.scatter(outliers['Date'], outliers['Stock Price'], color='red', label='Outliers')
plt.title('Stock Price with Outliers')
plt.xlabel('Date')
plt.ylabel('Stock Price')
plt.legend()
plt.show()

问题根源:滚动窗口的「污染」效应

你的代码核心问题在于启用了center=True的30天滚动窗口:

  • 当窗口覆盖到异常时间段时,异常值会直接拉高滚动均值,同时让滚动标准差急剧增大,导致上下边界跟着异常值一起抬高。
  • 原本的异常值落在了被污染后的边界范围内,自然无法被标记为异常。

另外,代码中df的Date列与shifted_prices的索引重复,属于冗余,但不影响异常检测逻辑。

修正方案

方案1:使用非中心滚动窗口(向前窗口)

将center=True改为center=False,让滚动窗口仅计算当前日期之前的30天数据,这样异常值出现时,前30天都是正常数据,边界不会被异常值污染,能立刻识别出异常:

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt

baseline_price = 100.00
dates = pd.date_range(start='2024-01-01', end='2024-12-31', freq='D')
baseline_prices = pd.Series(baseline_price, index=dates)
baseline_prices += np.random.uniform(10, 15, size=len(baseline_prices))

shifted_prices = baseline_prices.copy()
shifted_prices.loc['2024-04-01':'2024-05-30'] += np.random.uniform(1000, 1015, size=60)

# 直接用索引作为日期,避免冗余列
df = pd.DataFrame({'Stock Price': shifted_prices})

window_size, multiplier = 30, 2
# 使用非中心窗口,仅基于历史数据计算均值和标准差
rolling_mean = df['Stock Price'].rolling(window=window_size, center=False).mean()
rolling_std = df['Stock Price'].rolling(window=window_size, center=False).std()

upper_bound = rolling_mean + (multiplier * rolling_std)
lower_bound = rolling_mean - (multiplier * rolling_std)
outliers = df[(df['Stock Price'] > upper_bound) | (df['Stock Price'] < lower_bound)]

plt.figure(figsize=(10, 6))
plt.plot(df.index, df['Stock Price'], label='Stock Price')
plt.plot(df.index, rolling_mean, label=f'{window_size}-day Rolling Mean')
plt.fill_between(df.index, upper_bound, lower_bound, color='gray', alpha=0.2, label=f'{multiplier}-sigma Bounds')
plt.scatter(outliers.index, outliers['Stock Price'], color='red', label='Outliers')
plt.title('Stock Price with Outliers (Fixed: Non-centered Rolling Window)')
plt.xlabel('Date')
plt.ylabel('Stock Price')
plt.legend()
plt.show()

方案2:使用固定基线计算边界(适合已知正常基准场景)

如果数据有明确的正常基线(比如异常时间段之前的数据均为正常),可以直接用基线数据的均值和标准差计算固定边界,完全避免异常值的干扰:

# 仅用异常发生前的正常数据计算基准
baseline_data = baseline_prices.loc[:'2024-03-31']
baseline_mean = baseline_data.mean()
baseline_std = baseline_data.std()

# 生成固定边界
upper_bound = baseline_mean + multiplier * baseline_std
lower_bound = baseline_mean - multiplier * baseline_std

# 检测异常值
outliers = df[(df['Stock Price'] > upper_bound) | (df['Stock Price'] < lower_bound)]

内容的提问来源于stack exchange,提问作者r ram

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 09:46:23