You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于矩阵表达式$X(X^{T}X)^{-1}X^{T}$化简及相关期望推导的技术咨询

关于矩阵表达式$X(X{T}X){-1}X^{T}$化简及相关期望推导的技术咨询

Hey Tim, let's work through your confusion step by step—these are exactly the kinds of details that trip up even experienced folks when working with non-square matrices, so great job digging into the ambiguity!

First, let's unpack your two core problems:

  • Problem 1: Misapplying inverse rules to non-square matrices
    The rule $(AB)^{-1} = B{-1}A{-1}$ only applies when both A and B are invertible square matrices. Your X is $n \times k$ (non-square, assuming $n \neq k$), so $X^{-1}$ doesn't even exist—non-square matrices don't have standard inverses (they might have pseudoinverses, but that's a different tool). When you tried to expand $(XTX){-1}$ into $X{-1}(XT)^{-1}$, you broke this critical rule, which is why the dimension check failed: $(n \times k) \times (n \times k)$ isn't a valid matrix multiplication (inner dimensions don't match). That step was invalid from the start.

  • Problem 2: Why you can't get $I$ directly from $X(XTX){-1}X^T$
    Your second instinct was correct: $X^TX$ is a $k \times k$ square matrix (since $X^T$ is $k \times n$ and $X$ is $n \times k$), and if X has full column rank (a standard assumption in regression), $X^TX$ is invertible. But $X(XTX){-1}X^T$ isn't a product of a matrix and its inverse—it's a projection matrix (often written as $P$) that maps vectors onto the column space of X. This matrix is $n \times n$, and it only equals $I_n$ if $n = k$ (i.e., X is square and invertible). For $n > k$ (the typical regression setup, more observations than parameters), $P$ is not the identity matrix.

Your key question: What's the dimension of the "I" in the derivation?

The derivation you're verifying doesn't claim $X(XTX){-1}X^T = I$—it uses a trick with the trace operator to simplify the expression. Here's the breakdown of that critical step:

  1. By the cyclic property of the trace ($\text{trace}(AB) = \text{trace}(BA)$), we can rearrange the product inside the trace:
    $$\text{trace}\left(X(XTX){-1}X^T\right) = \text{trace}\left((XTX){-1}X^TX\right)$$
  2. Now, $(XTX){-1}X^TX$ is exactly the product of a $k \times k$ matrix and its inverse, so this equals $I_k$ (the $k \times k$ identity matrix).
  3. The trace of $I_k$ is just $k$, and the trace of $I_n$ is $n$, which gives you the final result $\sigma^2(n - k)$.

A quick walkthrough of your full derivation to confirm it's correct:

$$\begin{aligned}
\mathbf{E}\left(e^{\prime} e\right) & =\mathbf{E}\left(\operatorname{trace}\left(e^{\prime} e\right)\right)=\mathbf{E}\left(\operatorname{trace}\left(e e^{\prime}\right)\right)=\operatorname{trace}\left(\mathbf{E}\left(e e^{\prime}\right)\right) \
& =\operatorname{trace}\left(\sigma^2 M\right)=\sigma^2 \operatorname{trace}\left(I_n-X\left(X^{\prime} X\right)^{-1} X^{\prime}\right) \
& =\sigma^2 \operatorname{trace}\left(I_n\right)-\sigma^2 \operatorname{trace}\left(X\left(X^{\prime} X\right)^{-1} X^{\prime}\right) \
& =\sigma^2 \operatorname{trace}\left(I_n\right)-\sigma^2 \operatorname{trace}\left(\left(X^{\prime} X\right)^{-1} X^{\prime} X\right) \
& =\sigma^2 \operatorname{trace}\left(I_n\right)-\sigma^2 \operatorname{trace}\left(I_k\right) \
& =\sigma^2(n-k)
\end{aligned}$$

Every step checks out:

  • $\mathbf{E}(e'e) = \mathbf{E}(\text{trace}(e'e))$: A scalar's trace is itself, so this is trivial.
  • $\text{trace}(e'e) = \text{trace}(ee')$: Cyclic trace property (works even for non-square matrices, as long as both products are defined).
  • $\mathbf{E}(\text{trace}(ee')) = \text{trace}(\mathbf{E}(ee'))$: Trace is a linear operator, and expectation is linear, so they commute.
  • $\text{trace}(I_n - P) = \text{trace}(I_n) - \text{trace}(P)$: Another linearity property of the trace.
  • The cyclic trace trick to turn $\text{trace}(P)$ into $\text{trace}(I_k)$: As we explained earlier, this is the key move that makes the derivation work.

All in all, your confusion came from a common misapplication of inverse rules to non-square matrices—once you fix that and lean on trace properties instead, everything falls into place!

备注:内容来源于stack exchange,提问作者Tim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.23 13:37:35