多元二元时间序列平稳化与AR(p)模型构建技术问询
Great question! Let's walk through this step by step since you're working with a 10-dimensional binary time series and want to build an AR(p)-style model while handling stationarity and using link functions correctly.
Traditional stationarity techniques (like differencing for continuous series) don't apply here because your data is discrete binary. Instead, we focus on distributional stationarity: the joint probability distribution of your time series shouldn't shift over time (i.e., $P(Y_t, Y_{t-1}, ...) = P(Y_{t+k}, Y_{t+k-1}, ...)$ for any time lag $k$).
- If your series has a time trend (e.g., the probability of each binary variable being 1 drifts over time), first fit a simple model with time as a covariate to strip out this trend. Then you can focus on modeling the residual AR structure.
- For state-dependent non-stationarity, you might need a Markov-switching framework, but let's stick to your AR(p) goal first.
Your proposed formula $P(Y_t \mid Y_{t-1}, ... Y_{t-p}) = \ell(\beta_0 + \beta_1Y_{t-1} + ... + \beta_p Y_{t-p})$ is on the right track, but we need to clarify the structure for a 10-dimensional $Y_t$:
2.1 Link Functions: Why You Need Them
The linear combination $\beta_0 + \beta_1Y_{t-1} + ... + \beta_p Y_{t-p}$ outputs real numbers, but probabilities must lie between 0 and 1. The link function $\ell$ maps those real values to the [0,1] interval. Your top choices are:
- Logit Link: $\ell(x) = \frac{ex}{1+ex}$ — this is the most common, with intuitive odds ratio interpretations and easy computation.
- Probit Link: $\ell(x) = \Phi(x)$ (where $\Phi$ is the standard normal CDF) — ideal if you want to model underlying latent normal variables.
- Cloglog Link: $\ell(x) = 1 - e{-ex}$ — useful for right-skewed response probabilities.
2.2 Model Structure (Two Main Options)
Since $Y_t$ is 10-dimensional, you have two approaches to model the conditional probability:
Option 1: Component-Wise AR(p) with Cross-Dependencies
Treat each binary variable $Y_{t,j}$ (for $j=1$ to 10) as a separate response, but include lags of all 10 variables as predictors. For each component, the formula becomes:
$$
P(Y_{t,j}=1 \mid Y_{t-1}, ..., Y_{t-p}) = \ell_j\left( \beta_{0,j} + \sum_{k=1}^p \sum_{m=1}^{10} \beta_{k,j,m} Y_{t-k,m} \right)
$$
- $\beta_{0,j}$ is the intercept for the $j$-th component.
- $\beta_{k,j,m}$ measures how the $m$-th variable at lag $k$ affects the probability that the $j$-th variable is 1 at time $t$.
- You can use the same link function $\ell$ for all components or pick different ones based on each variable's behavior.
Option 2: Joint Multivariate AR(p) Model
If you want to model the joint distribution of all 10 variables at time $t$ (accounting for correlations between them), you'll need a multivariate link function. For example:
- A multivariate probit model uses latent normal variables with a covariance matrix to capture cross-component correlations.
- A copula-based model fits individual AR(p) models for each component, then uses a copula to model their joint dependence.
This is more complex, but necessary if the binary variables aren't conditionally independent given the lagged values.
2.3 Deriving the Full Conditional Probability
If you assume conditional independence between the 10 components (given lags), the joint conditional probability is the product of each component's conditional probability:
$$
P(Y_t \mid Y_{t-1}, ..., Y_{t-p}) = \prod_{j=1}^{10} \left[ P(Y_{t,j}=1 \mid \text{lags}) \right]^{Y_{t,j}} \left[ 1 - P(Y_{t,j}=1 \mid \text{lags}) \right]^{1-Y_{t,j}}
$$
For a joint model (no conditional independence), the formula will involve integrating over latent variables (for probit) or using copula density functions.
- Lag Selection: Use AIC or BIC to choose the optimal $p$, or use cross-validation. Keep in mind that $p$ increases the number of parameters exponentially (10 components * (1 + 10*p) parameters), so avoid overfitting.
- Stationarity Validation: For non-linear models like this, check if the coefficient matrices have eigenvalues with magnitude less than 1 (similar to linear multivariate AR models). You can also simulate from the fitted model to see if the simulated series matches the stationarity of your data.
- Estimation: In R, use
glmfor component-wise logit/probit models, ormvtnormfor multivariate probit. In Python,statsmodelshasLogitfor single components, andpymc3can handle Bayesian estimation for more complex multivariate models.
To make this concrete, here's the formula for the $j$-th component with $p=1$ and a logit link:
$$
P(Y_{t,j}=1 \mid Y_{t-1}) = \frac{e^{\beta_{0,j} + \beta_{1,j,1}Y_{t-1,1} + ... + \beta_{1,j,10}Y_{t-1,10}}}{1 + e^{\beta_{0,j} + \beta_{1,j,1}Y_{t-1,1} + ... + \beta_{1,j,10}Y_{t-1,10}}}
$$
Here, $\beta_{1,j,m}$ tells you how much the log-odds of $Y_{t,j}=1$ increases when $Y_{t-1,m}=1$.
内容的提问来源于stack exchange,提问作者StatsSorceress

