卡尔曼滤波器证明中的矩阵方程梯度求解疑问
Hey there, let's walk through how to compute the gradient of your cost function step by step—since you're working through Kalman filter proofs, I know getting these matrix calculus details right is crucial.
First, let's restate your cost function clearly for reference:
$$
J(x) = \frac{1}{2}(y - \mathcal{H}(x))^T R^{-1} (y - \mathcal{H}(x))
$$
With the given definitions:
- $x \in \mathbb{R}^n$ (the variable we're optimizing over)
- $y \in \mathbb{R}^m$ (constant measurement vector)
- $\mathcal{H}: \mathbb{R}^n \rightarrow \mathbb{R}^m$ (nonlinear measurement function, in general)
- $R \in \mathbb{R}^{m \times m}$ (symmetric positive-definite noise covariance matrix, so $R^{-1}$ is also symmetric)
Step 1: Simplify with an Error Term
Let's define the measurement error vector to make the math cleaner:
$$
e(x) = y - \mathcal{H}(x)
$$
This rewrites the cost function as:
$$
J(x) = \frac{1}{2} e(x)^T R^{-1} e(x)
$$
Step 2: Use Matrix Calculus & Chain Rule
We need to compute $\nabla_x J(x)$—the gradient of the scalar $J(x)$ with respect to the vector $x$, which will be an $n$-dimensional vector.
Key Matrix Calculus Rule
For a scalar function of the form $J = \frac{1}{2} v^T A v$ where $A$ is symmetric and $v$ is a vector, the gradient with respect to $v$ is:
$$
\nabla_v J = A v
$$
We'll use this, plus the chain rule for vector-valued functions.
Apply the Chain Rule
First, compute the gradient of $J$ with respect to $e(x)$:
$$
\nabla_e J = R^{-1} e(x)
$$
(This comes directly from the rule above, since $R^{-1}$ is symmetric.)Next, compute the Jacobian matrix of $e(x)$ with respect to $x$. Since $e(x) = y - \mathcal{H}(x)$, the Jacobian (denoted $H(x)$) is the negative of $\mathcal{H}(x)$'s Jacobian:
$$
H(x) = \nabla_x \mathcal{H}(x) = \begin{bmatrix}
\frac{\partial \mathcal{H}_1}{\partial x_1} & \dots & \frac{\partial \mathcal{H}_1}{\partial x_n} \
\vdots & \ddots & \vdots \
\frac{\partial \mathcal{H}_m}{\partial x_1} & \dots & \frac{\partial \mathcal{H}_m}{\partial x_n}
\end{bmatrix}
$$
So $\nabla_x e(x) = -H(x)$ (this is an $m \times n$ matrix).Finally, chain these together. To confirm, let's do a component-wise check (it's easier to wrap your head around):
Take any component $x_k$ of $x$. Compute $\frac{\partial J}{\partial x_k}$:
$$
\frac{\partial J}{\partial x_k} = \frac{1}{2} \left( \frac{\partial e^T}{\partial x_k} R^{-1} e + e^T R^{-1} \frac{\partial e}{\partial x_k} \right)
$$
Since $R^{-1}$ is symmetric, $e^T R^{-1} \frac{\partial e}{\partial x_k} = \left( \frac{\partial e}{\partial x_k} \right)^T R^{-1} e$. So the two terms add up:
$$
\frac{\partial J}{\partial x_k} = \left( \frac{\partial e}{\partial x_k} \right)^T R^{-1} e
$$
Substitute $\frac{\partial e}{\partial x_k} = -H(:,k)$ (the $k$-th column of $H(x)$):
$$
\frac{\partial J}{\partial x_k} = -H(:,k)^T R^{-1} (y - \mathcal{H}(x))
$$
Final Gradient Result
Stacking all these components into a vector gives the full gradient:
$$
\nabla_x J(x) = -H(x)^T R^{-1} (y - \mathcal{H}(x))
$$
Note for First-Order Approximation
When doing a first-order (linear) approximation of $\mathcal{H}(x)$ around a reference point $x_0$, we write:
$$
\mathcal{H}(x) \approx \mathcal{H}(x_0) + H(x_0)(x - x_0)
$$
Substituting this into the gradient expression just replaces $H(x)$ with $H(x_0)$ (the Jacobian evaluated at $x_0$)—this is exactly the linearized gradient used in the Extended Kalman Filter (EKF) update step.
内容的提问来源于stack exchange,提问作者Sebastian Garrido

