多时间序列符号回归:寻求多变量对目标时序变量的影响分析方法
Got it, let's break down how to analyze the impact of columns a and b on your target time series c. You're right that pandas' corr() only gives pairwise correlations, and ARMA/ARIMA are univariate—so here are the best methods to solve your problem, including symbolic regression and other multi-variable time-series tools:
1. Symbolic Regression (Your Exact Request)
Symbolic regression is perfect here because it doesn't force you to predefine a model structure (like linear regression does). Instead, it evolves mathematical expressions (using operators like +, -, *, /, sqrt, log) that map your input variables (a, b, and their lagged values) to your target c.
- Tool to use: The
gplearnlibrary in Python is a great open-source option. Here's a quick snippet to get you started:from gplearn.genetic import SymbolicRegressor import pandas as pd # Create lag features for a and b (e.g., lag 1 to 3) df = pd.DataFrame({'a': your_a_data, 'b': your_b_data, 'c': your_c_data}) for lag in range(1, 4): df[f'a_lag_{lag}'] = df['a'].shift(lag) df[f'b_lag_{lag}'] = df['b'].shift(lag) df = df.dropna() # Define features (a, b, their lags) and target (c) X = df[['a', 'b', 'a_lag_1', 'b_lag_1', 'a_lag_2', 'b_lag_2', 'a_lag_3', 'b_lag_3']] y = df['c'] # Train symbolic regressor sr = SymbolicRegressor(population_size=5000, generations=20, verbose=1) sr.fit(X, y) # Print the best found expression linking a/b to c print(sr._program) - Pros: Discovers non-obvious, non-linear relationships between a/b and c that you might miss with traditional models.
- Cons: Can be computationally expensive; results might be hard to interpret if the evolved expression gets too complex.
2. Multivariate Autoregressive Models (VAR/VARMA/VARIMA)
Quick correction: while ARMA/ARIMA are univariate, their multivariate counterparts are built explicitly to model interactions between multiple time series. These models assume that your target c depends on its own past values and the past values of a and b.
- Tool to use: The
statsmodelslibrary has solid support for VAR/VARMA. Example workflow:from statsmodels.tsa.vector_ar.var_model import VAR # Prepare multivariate time series (a, b, c) data = df[['a', 'b', 'c']].dropna() # Fit VAR model (choose lag order based on AIC/BIC scores) model = VAR(data) results = model.fit(maxlags=3) # View coefficients for c's equation (shows how lags of a/b impact c) print("Coefficients for c's equation:") print(results.coefs[:, 2]) # Column 2 corresponds to the target c - Pros: Highly interpretable; ideal for linear or weakly non-linear relationships.
- Cons: Struggles with complex non-linear patterns that symbolic regression might catch.
3. Non-Linear Time Series Models (ML-Based)
If your data has strong non-linear dependencies, machine learning models can capture them while letting you quantify the impact of a and b on c:
- Tree-Based Models (XGBoost/LightGBM): Convert your time series into supervised learning data by adding lag features for a, b, and c. Train a regressor, then use feature importance scores to compare how much a (and its lags) vs. b (and its lags) contribute to predicting c.
- Deep Learning (LSTM/GRU): Use recurrent neural networks to model sequential relationships. You can add attention mechanisms or calculate permutation importance to infer the relative impact of a and b on c.
4. Granger Causality (For Causal Insight)
While not a regression method, Granger causality tests complement your analysis by checking whether past values of a or b help predict c (beyond using past values of c alone). This is a quick way to confirm if a/b have a meaningful predictive relationship with c before diving into regression.
- Example with statsmodels:
from statsmodels.tsa.stattools import grangercausalitytests # Test if a Granger-causes c grangercausalitytests(df[['c', 'a']], maxlag=3) # Test if b Granger-causes c grangercausalitytests(df[['c', 'b']], maxlag=3)
内容的提问来源于stack exchange,提问作者tardis

