递归最小二乘算法变量咨询:K(t)、φ(t)、P(t)具体指代什么?
Hey there, let’s unpack each of those Recursive Least Squares (RLS) variables with concrete, intuitive explanations—no more vague "matrix/vector" hand-waving!
First, let’s ground this in context: RLS is an online algorithm for estimating parameters of a linear model at each time step ( t ). The core model we’re fitting looks like this:
$$y(t) = \phi^T(t)\theta + \epsilon(t)$$
where ( \epsilon(t) ) is random noise in our observed output ( y(t) ). Now let’s break down each variable you’re curious about:
( \phi(t) ): The Regressor (Feature) Vector
This is your input/feature vector at time ( t )—it’s the set of variables you believe directly drive the observed output ( y(t) ).
For example:
- If you’re fitting a simple linear model ( y(t) = a \cdot x(t) + b ), then ( \phi(t) = \begin{bmatrix} x(t) \ 1 \end{bmatrix} ) (the ( 1 ) captures the intercept term ( b )).
- If you’re modeling a system with 3 input sensors, ( \phi(t) = \begin{bmatrix} \text{sensor1}(t) \ \text{sensor2}(t) \ \text{sensor3}(t) \end{bmatrix} ).
Its dimension matches the parameter vector ( \hat{\theta}(t) )—each element corresponds to a parameter you’re trying to estimate.
( K(t) ): The Gain Matrix
Think of this as your update weight matrix. It tells you how much to adjust your old parameter estimate ( \hat{\theta}(t-1) ) based on the current prediction error ( (y(t) - \phi^T(t)\hat{\theta}(t-1)) ).
- If ( K(t) ) has large values, we’re trusting the new observation ( y(t) ) more, so we make a big adjustment to ( \hat{\theta}(t-1) ).
- If ( K(t) ) is small, we’re leaning more on our past estimates, so the update is tiny.
The formula ( K(t) = P(t)\phi(t) ) ties it directly to our confidence in the current parameters (via ( P(t) )) and the current input features ( \phi(t) ).
( P(t) ): The Parameter Uncertainty Matrix
This is the inverse of the covariance matrix of your parameter estimation error (loosely, it’s a measure of how "confident" we are in our current parameter estimate ( \hat{\theta}(t) )).
- Small values in ( P(t) ) mean we’re very confident in the corresponding parameter—we don’t want to update it much.
- Large values mean the parameter estimate is uncertain, so we’re willing to make bigger updates based on new data.
The initial condition ( P(t_0) ) is usually set to a large diagonal matrix (e.g., ( 10^6 \cdot I )) because at the start, we have no prior information about the parameters, so we allow large adjustments as we gather data. The update rule ( P(t) = (I-K(t)\phi^T(t))P(t-1) ) refines our confidence over time: as we get more data, ( P(t) ) typically shrinks, reducing future update sizes.
Quick Concrete Example
Let’s take a single-variable linear model: ( y(t) = \theta \cdot x(t) + \epsilon(t) ) (so ( \theta ) is a scalar). The formulas simplify to:
- ( \phi(t) = x(t) ) (a scalar input)
- ( K(t) = P(t) \cdot x(t) )
- ( \hat{\theta}(t) = \hat{\theta}(t-1) + K(t) \cdot (y(t) - x(t)\hat{\theta}(t-1)) )
- ( P(t) = (1 - K(t) \cdot x(t)) \cdot P(t-1) )
If you start with ( \hat{\theta}(0) = 0 ) and ( P(0) = 1000 ) (high uncertainty), the first time you get ( x(1)=2 ) and ( y(1)=5 ):
- ( K(1) = 1000 \cdot 2 = 2000 ) (big gain because we’re uncertain)
- ( \hat{\theta}(1) = 0 + 2000 \cdot (5 - 2\cdot0) = 10000 ) (huge initial update)
- ( P(1) = (1 - 2000\cdot2)\cdot1000 = -3999000 ) (note: this is a quirk of the rearranged formula you shared; the standard scalar RLS gain avoids negatives, but the core intuition holds: early gains are large, then shrink as ( P(t) ) gets smaller with more data.)
Hope this makes each variable feel tangible instead of just abstract math!
内容的提问来源于stack exchange,提问作者euraad

