时序数据异常值标记失效:逻辑问题还是代码问题?
时序数据异常值检测问题:为何添加的异常未被识别?
尝试用Pandas简单检测时序数据异常值,在2024-04-01至2024-05-30时间段添加了大幅噪声,但代码未检测到这些异常值,请问是逻辑缺陷还是代码错误?
import pandas as pd import numpy as np import matplotlib.pyplot as plt baseline_price = 100.00 dates = pd.date_range(start='2024-01-01', end='2024-12-31', freq='D') baseline_prices = pd.Series(baseline_price, index=dates) baseline_prices += np.random.uniform(10, 15, size=len(baseline_prices)) shifted_prices = baseline_prices.copy() shifted_prices.loc['2024-04-01':'2024-05-30'] += np.random.uniform(1000, 1015, size=60) df = pd.DataFrame({'Date': dates, 'Stock Price': shifted_prices}) window_size, multiplier = 30, 2 rolling_mean = df['Stock Price'].rolling(window=window_size, center=True).mean() rolling_std = df['Stock Price'].rolling(window=window_size, center=True).std() upper_bound = rolling_mean + (multiplier * rolling_std) lower_bound = rolling_mean - (multiplier * rolling_std) outliers = df[(df['Stock Price'] > upper_bound) | (df['Stock Price'] < lower_bound)] plt.figure(figsize=(10, 6)) plt.plot(df['Date'], df['Stock Price'], label='Stock Price') plt.plot(df['Date'], rolling_mean, label=f'{window_size}-day Rolling Mean') plt.fill_between(df['Date'], upper_bound, lower_bound, color='gray', alpha=0.2, label=f'{multiplier}-sigma Bounds') plt.scatter(outliers['Date'], outliers['Stock Price'], color='red', label='Outliers') plt.title('Stock Price with Outliers') plt.xlabel('Date') plt.ylabel('Stock Price') plt.legend() plt.show()
问题根源:滚动窗口的「污染」效应
你的代码核心问题在于启用了center=True的30天滚动窗口:
- 当窗口覆盖到异常时间段时,异常值会直接拉高滚动均值,同时让滚动标准差急剧增大,导致上下边界跟着异常值一起抬高。
- 原本的异常值落在了被污染后的边界范围内,自然无法被标记为异常。
另外,代码中df的Date列与shifted_prices的索引重复,属于冗余,但不影响异常检测逻辑。
修正方案
方案1:使用非中心滚动窗口(向前窗口)
将center=True改为center=False,让滚动窗口仅计算当前日期之前的30天数据,这样异常值出现时,前30天都是正常数据,边界不会被异常值污染,能立刻识别出异常:
import pandas as pd import numpy as np import matplotlib.pyplot as plt baseline_price = 100.00 dates = pd.date_range(start='2024-01-01', end='2024-12-31', freq='D') baseline_prices = pd.Series(baseline_price, index=dates) baseline_prices += np.random.uniform(10, 15, size=len(baseline_prices)) shifted_prices = baseline_prices.copy() shifted_prices.loc['2024-04-01':'2024-05-30'] += np.random.uniform(1000, 1015, size=60) # 直接用索引作为日期,避免冗余列 df = pd.DataFrame({'Stock Price': shifted_prices}) window_size, multiplier = 30, 2 # 使用非中心窗口,仅基于历史数据计算均值和标准差 rolling_mean = df['Stock Price'].rolling(window=window_size, center=False).mean() rolling_std = df['Stock Price'].rolling(window=window_size, center=False).std() upper_bound = rolling_mean + (multiplier * rolling_std) lower_bound = rolling_mean - (multiplier * rolling_std) outliers = df[(df['Stock Price'] > upper_bound) | (df['Stock Price'] < lower_bound)] plt.figure(figsize=(10, 6)) plt.plot(df.index, df['Stock Price'], label='Stock Price') plt.plot(df.index, rolling_mean, label=f'{window_size}-day Rolling Mean') plt.fill_between(df.index, upper_bound, lower_bound, color='gray', alpha=0.2, label=f'{multiplier}-sigma Bounds') plt.scatter(outliers.index, outliers['Stock Price'], color='red', label='Outliers') plt.title('Stock Price with Outliers (Fixed: Non-centered Rolling Window)') plt.xlabel('Date') plt.ylabel('Stock Price') plt.legend() plt.show()
方案2:使用固定基线计算边界(适合已知正常基准场景)
如果数据有明确的正常基线(比如异常时间段之前的数据均为正常),可以直接用基线数据的均值和标准差计算固定边界,完全避免异常值的干扰:
# 仅用异常发生前的正常数据计算基准 baseline_data = baseline_prices.loc[:'2024-03-31'] baseline_mean = baseline_data.mean() baseline_std = baseline_data.std() # 生成固定边界 upper_bound = baseline_mean + multiplier * baseline_std lower_bound = baseline_mean - multiplier * baseline_std # 检测异常值 outliers = df[(df['Stock Price'] > upper_bound) | (df['Stock Price'] < lower_bound)]
内容的提问来源于stack exchange,提问作者r ram
相关产品推荐
相关产品推荐

