带最优滞后阶数的VAR多变量预测:72列数据批量处理需求
Got it, let's tackle this problem step by step. Since you need to handle 72 endogenous variables (like dj1 to dj72) and generate point predictions for each, here's a practical, scalable approach tailored to your needs:
Step 1: Automate Endogenous Variable List
First, skip manually typing 72 variable names—generate the list programmatically:
# Create a list of your 72 endogenous variable names (dj1 to dj72) endog_vars = [f"dj{i}" for i in range(1, 73)]
Step 2: Initialize a Results Container
We'll use a pandas DataFrame to store all predictions neatly:
import pandas as pd import statsmodels.api as sm # Adjust based on your preferred model library
Step 3: Loop Through Each Endogenous Variable
For each dj variable, fit your model, generate point predictions, and store the results. Below is a generic example using OLS (swap in your actual model type—ARIMA, GLM, etc.—as needed):
# Initialize empty DataFrame to hold predictions (matches your original data's index) predictions_df = pd.DataFrame(index=dataframe.index) for var in endog_vars: # Define your model's endogenous (y) and exogenous (X) variables y = dataframe[var] # Example: Use all non-dj columns as exogenous data (adjust to your actual equation spec!) X = dataframe.drop(columns=endog_vars) X = sm.add_constant(X) # Add intercept if your model requires it # Fit the model model = sm.OLS(y, X).fit() # Generate point predictions (in-sample shown; use out-of-sample data if needed) var_predictions = model.predict(X) # Add predictions to our results DataFrame predictions_df[var] = var_predictions
Step 4: Get Predictions as Matrix (If Needed)
If you need a matrix instead of a DataFrame, just convert it:
predictions_matrix = predictions_df.values
Bonus: For Multi-Variable Models (Like VAR)
If your 72 dj variables are modeled together (e.g., Vector Autoregression), skip the loop and fit a single multi-variable model for efficiency:
from statsmodels.tsa.vector_ar.var_model import VAR # Subset your data to only the endogenous variables endog_data = dataframe[endog_vars] # Fit the VAR model (adjust maxlags to your needs) var_model = VAR(endog_data) fitted_var = var_model.fit(maxlags=2) # Generate point predictions (e.g., 1 step ahead; adjust steps as needed) var_predictions = fitted_var.forecast(fitted_var.y, steps=1) # Convert to DataFrame for readability predictions_df = pd.DataFrame(var_predictions, columns=endog_vars)
Key Notes
- Swap out the model fitting code (
sm.OLS,VAR, etc.) with whatever model you're actually using for your equations. - Adjust the exogenous variable selection (
X = dataframe.drop(...)) to match your specific equation specifications. - For out-of-sample predictions, pass your out-of-sample exogenous data to the
predict()method instead of the in-sampleX.
内容的提问来源于stack exchange,提问作者DaWassi

