时间序列入门:如何用移动平均/滚动均值及指数平滑预测t+1值?
Hey there! Let's break this down step by step since you're just getting started with time series forecasting—no fancy stuff needed here, just solid basics to get you predicting that 2018 Volume data. I'll cover both moving averages and exponential smoothing, including how to calculate t+1 predictions and visualize them alongside your existing data.
First, let's recap: moving averages smooth out noise by averaging the last N periods of data. To predict the next period (t+1), you just take the average of the most recent N periods you have. Here's how to implement this with your pandas DataFrame:
Step-by-Step Code
import pandas as pd import matplotlib.pyplot as plt # First, make sure your date column is in datetime format (critical for time series!) df['date'] = pd.to_datetime(df['date']) df.set_index('date', inplace=True) # Choose your window size (e.g., 3 periods for short-term, 12 for monthly annual trends) window_size = 3 df[f'MA_{window_size}'] = df['Volume'].rolling(window=window_size).mean() # Calculate t+1 prediction: average of the last N periods t_plus_1_pred_ma = df['Volume'].tail(window_size).mean() # Add the prediction to your DataFrame (replace the date with your actual next period) next_period = pd.to_datetime('2018-01-01') # Adjust this to match your data's frequency df.loc[next_period, 'MA_Prediction'] = t_plus_1_pred_ma # Visualize everything together plt.figure(figsize=(10, 6)) plt.plot(df['Volume'], label='Actual Volume') plt.plot(df[f'MA_{window_size}'], label=f'{window_size}-Period Moving Average') plt.scatter(next_period, t_plus_1_pred_ma, color='red', s=100, label='t+1 Prediction') plt.title('Actual Volume vs. Moving Average + Prediction') plt.xlabel('Date') plt.ylabel('Volume') plt.legend() plt.show()
Quick Notes
- Adjust
window_sizebased on your data frequency: use 12 for monthly data (annual cycle), 7 for daily (weekly cycle), etc. - If you need to predict the entire 2018 year (not just one period), repeat this logic: each new prediction uses the previous N values (including prior predictions if forecasting multiple steps ahead, though accuracy drops further out).
Exponential smoothing is better than simple moving averages because it gives more weight to recent data (instead of treating all N periods equally). For your low-precision needs, Simple Exponential Smoothing (SES) is perfect—it works best for data with no obvious trend or seasonality.
Step-by-Step Code
We'll use the statsmodels library (built for time series and super easy to use):
from statsmodels.tsa.holtwinters import SimpleExpSmoothing # Fit the simple exponential smoothing model # The smoothing_level (α) controls how much weight we give recent data (0 < α < 1) # Leave it blank to let the model automatically optimize α for your data model = SimpleExpSmoothing(df['Volume']) fit_model = model.fit(smoothing_level=0.3) # α=0.3: moderate weight on recent data # Predict t+1 value t_plus_1_pred_ses = fit_model.predict(start=len(df), end=len(df)).iloc[0] # Add the prediction to your DataFrame df.loc[next_period, 'SES_Prediction'] = t_plus_1_pred_ses # Visualize the fitted values and prediction plt.figure(figsize=(10, 6)) plt.plot(df['Volume'], label='Actual Volume') plt.plot(fit_model.fittedvalues, label='Fitted Exponential Smoothing') plt.scatter(next_period, t_plus_1_pred_ses, color='green', s=100, label='t+1 SES Prediction') plt.title('Actual Volume vs. Exponential Smoothing + Prediction') plt.xlabel('Date') plt.ylabel('Volume') plt.legend() plt.show() # Bonus: Predict the entire 2018 year (e.g., 12 months for monthly data) future_periods = 12 # Adjust based on how many periods you need to predict future_dates = pd.date_range(start=next_period, periods=future_periods, freq='M') # 'M' for monthly future_preds = fit_model.predict(start=len(df), end=len(df) + future_periods - 1) # Combine with original data for visualization future_df = pd.DataFrame({'SES_Annual_Prediction': future_preds}, index=future_dates) combined_df = pd.concat([df, future_df]) plt.figure(figsize=(12, 6)) plt.plot(combined_df['Volume'], label='Actual Volume') plt.plot(combined_df['SES_Annual_Prediction'], label='2018 Volume Predictions', color='green', linestyle='--') plt.title('Actual Volume + 2018 Exponential Smoothing Predictions') plt.xlabel('Date') plt.ylabel('Volume') plt.legend() plt.show()
Quick Notes
- If your data has a clear upward/downward trend, you could use Holt's Linear Trend (a slight extension of SES), but SES is more than enough for your low-precision needs.
- Letting the model optimize
smoothing_level(by removing the parameter fromfit()) will usually give better results than picking a value manually.
内容的提问来源于stack exchange,提问作者Daniel Chepenko

