在sktime中是否支持同时使用过去协变量与未来协变量?如何区分两类协变量?
1. Can sktime handle both past and future covariates at the same time?
Absolutely! Sktime is built to natively support this use case—combining historical observed data with known future information is super common in real-world forecasting (think weather forecasts for retail sales, or scheduled promotions for inventory planning), and the library’s API is designed to accommodate this seamlessly.
2. How to distinguish between these two types of covariates?
The line between them boils down to when you pass the data to the model and how it aligns with your training/prediction timeframes:
Past covariates (historical/lagged covariates)
These are variables you have full observations for during your training period, and you want to use their lagged values as features to predict your target. For example, using day N-1’s observed temperature to forecast day N’s sales.
You pass these to thefit()method via theXparameter, making sure their time index matches exactly with your training targety. Sktime’s built-in tools (likeLagFeatures) can automatically generate lagged versions of these covariates for your model if needed.Future covariates (known future exogenous variables)
These are variables where you already have confirmed values for the entire prediction horizon before you run your forecast. For example, the day N weather forecast that’s available on day N-1.
You pass these to thepredict()method via theXparameter, ensuring their time index perfectly matches yourForecastingHorizon—every time point in your forecast needs a corresponding covariate value.
Example Code Walkthrough
Let’s use your sample data to show this in action:
import pandas as pd from sktime.forecasting.naive import NaiveForecaster from sktime.forecasting.base import ForecastingHorizon # Recreate your sample dataset df = pd.DataFrame({ "observed_temperature": [35, 28], "temperature_forecast": [37, 31], "sales": [100, 120] }, index=pd.to_datetime(["2021-01-01", "2021-01-02"])) # Step 1: Define your prediction horizon (predict 2021-01-02, made on 2021-01-01) fh = ForecastingHorizon(pd.to_datetime(["2021-01-02"]), is_relative=False) # Step 2: Prepare training data and past covariates y_train = df["sales"].iloc[:-1] # Sales data from 2021-01-01 past_covariates = df[["observed_temperature"]].iloc[:-1] # Observed temp from 2021-01-01 # Step 3: Fit the model with past covariates forecaster = NaiveForecaster(strategy="last", sp=1) forecaster.fit(y=y_train, X=past_covariates) # Step 4: Prepare future covariates (known temp forecast for 2021-01-02) future_covariates = df[["temperature_forecast"]].iloc[1:] # Temp forecast for the prediction date # Step 5: Generate the forecast using future covariates y_pred = forecaster.predict(fh=fh, X=future_covariates)
Most sktime forecasters (like AutoARIMA, ProphetWrapper, or GradientBoostingForecaster) follow this same pattern. Just double-check that your covariate data’s time indices are properly aligned with your training target and forecast horizon, and you’re good to go.
内容的提问来源于stack exchange,提问作者atsangarides

