statsmodels中AR参数取反的技术问询:约定差异与零滞后参数问题
signal.lfilter and Statistical ARMA Conventions Great questions—this is a super common point of confusion when bridging statistical time series modeling and signal processing implementations. Let’s unpack each question clearly:
1. Which convention requires AR parameters to be negated?
The sign flip comes from differences in how signal processing filters and statistical ARMA models write their core equations:
- In statistical ARMA, an AR(p) model is defined as a generative process for the time series:
( x_t = \phi_1 x_{t-1} + \phi_2 x_{t-2} + ... + \phi_p x_{t-p} + \epsilon_t )
Here, ( \phi_1 ... \phi_p ) are the AR coefficients you’d estimate from data (e.g., via ARIMA tools), representing the positive contribution of past observations to the current value. - In signal processing, tools like
scipy.signal.lfilteruse the standard causal linear difference equation convention for recursive filters, which rearranges terms to group all output (filtered signal) terms on the left:( a_0 y[n] + a_1 y[n-1] + ... + a_p y[n-p] = b_0 x[n] + b_1 x[n-1] + ... + b_q x[n-q] )
When simulating an AR process withlfilter, your input ( x[n] ) is the white noise ( \epsilon_t ), and the output ( y[n] ) is the AR time series. To match the statistical AR model, we rearrange the statistical equation to fit the filter form:
( y[n] - \phi_1 y[n-1] - ... - \phi_p y[n-p] = \epsilon[n] )
Comparing this to the filter equation, you’ll see that ( a_0 = 1 ), and ( a_1 = -\phi_1 ), ( a_2 = -\phi_2 ), etc. So the AR coefficients from statistics need to be negated to fitlfilter’s parameter convention.
2. If statistical fields don’t have this convention, how do we reconcile this?
You’re absolutely right—statistical ARMA models do not use negated coefficients. The confusion arises because we’re translating between two different frameworks with distinct goals:
- Statistical ARMA focuses on modeling how past observations predict the current value, so coefficients are written to directly represent the magnitude and direction of that predictive relationship.
- Signal processing filters are designed to transform input signals (e.g., denoise, smooth), so their equations are structured around defining the system’s recursive behavior, which naturally groups output terms on the left with positive coefficients for the current output and negative coefficients for lagged outputs (relative to the statistical convention).
This isn’t a conflict in conventions—it’s just two different ways to write the same underlying math, tailored to each field’s needs.
3. Why isn’t the zero-lag parameter ( =1 ) negated?
The zero-lag parameter in question is the coefficient for the current output term (( y[n] ) in the filter equation, or ( x_t ) in the statistical model). In both frameworks, this coefficient is 1 by default:
- In the statistical AR model, the current term ( x_t ) has an implicit coefficient of 1 (it’s the value we’re predicting).
- In
lfilter’s equation, ( a_0 ) (the zero-lag output coefficient) is set to 1 (if you pass anaarray without specifying it, the function normalizes to make ( a_0 = 1 )). Since this coefficient is identical in both conventions, there’s no need to negate it. The sign flip only applies to the lagged output coefficients (( a_1 ... a_p )), which correspond to the statistical AR coefficients ( \phi_1 ... \phi_p ).
Quick Example
Let’s say you have an AR(1) model with statistical coefficient ( \phi_1 = 0.7 ):
- Statistical equation: ( x_t = 0.7 x_{t-1} + \epsilon_t )
- To simulate this with
lfilter, you’d use:import numpy as np from scipy.signal import lfilter # White noise input eps = np.random.normal(0, 1, 1000) # lfilter parameters: b is [1] (input coefficient), a is [1, -0.7] (output coefficients) ar_series = lfilter(b=[1], a=[1, -0.7], x=eps)
Here, we negated the statistical AR coefficient 0.7 to get -0.7 for the lagged output term in a, while keeping the zero-lag term (1) unchanged.
内容的提问来源于stack exchange,提问作者dayum

