基于Pandas构建Sklearn时间序列预测的多步滞后特征矩阵X
Hey there! Creating those lag features for your Sklearn model is straightforward with Pandas' built-in shift() method. Let's walk through the steps to get your DataFrame into the exact format you want:
Step 1: Prepare Your Data
First, make sure your DateTime column is properly parsed as datetime type, and your data is sorted chronologically (this is critical for time series—you don't want shifted values pointing to the wrong time points!):
import pandas as pd # Assuming your raw DataFrame is named 'df' df['DateTime'] = pd.to_datetime(df['DateTime']) df = df.sort_values('DateTime').reset_index(drop=True)
Step 2: Align Column Names (Matches Your Target Format)
Rename the original "Close Price" column to "ClosePrice" to match your desired output structure:
df.rename(columns={'Close Price': 'ClosePrice'}, inplace=True)
Step 3: Generate Lag Features
Use Pandas' shift() method to create columns for the previous 3 close prices. Each shift(n) moves the column data down by n rows, so shift(1) gives the immediately preceding value, shift(2) gives the value two rows back, etc.:
df['ClosePrice-1'] = df['ClosePrice'].shift(1) df['ClosePrice-2'] = df['ClosePrice'].shift(2) df['ClosePrice-3'] = df['ClosePrice'].shift(3)
Step 4: Clean Missing Values
The first 3 rows will have NaN values in the lag columns (since there's no prior data to fill them). Drop these rows to get a clean dataset ready for modeling:
df = df.dropna().reset_index(drop=True)
Step 5: Reorder Columns (Optional)
If you want columns in the exact order you showed, reorder them like this:
df = df[['DateTime', 'ClosePrice', 'ClosePrice-1', 'ClosePrice-2', 'ClosePrice-3']]
Final Result
After running these steps, your DataFrame will look exactly like your target format:
| DateTime | ClosePrice | ClosePrice-1 | ClosePrice-2 | ClosePrice-3 |
|---|---|---|---|---|
| 2017-01-03 00:30:00 | 1.04662 | 1.04672 | 1.04575 | 1.04630 |
| 2017-01-03 01:00:00 | 1.04766 | 1.04662 | 1.04672 | 1.04575 |
Bonus: Automate for More Lags
If you ever need more than 3 lag features, you can use a loop to avoid repetitive code:
num_lags = 3 for n in range(1, num_lags + 1): df[f'ClosePrice-{n}'] = df['ClosePrice'].shift(n)
That's all! This gives you the feature matrix X you need for Sklearn, where each row combines the current close price with its three most recent historical values.
内容的提问来源于stack exchange,提问作者Hamzeh

