拉格朗日乘数法:为何目标函数梯度是约束梯度的线性组合?
Great question—let’s unpack this rigorously, starting from the geometry of constrained optimization and working our way to the formal calculus condition.
First, Let’s Restate the Problem
We’re targeting a local maximum of $f(x)$ under equality constraints, formalized as:
$$
\begin{align}
\max:, f(x),& \enspace x \in \mathbb{R}^{n} \
\text{s.t.} \quad g_{1}(x) &= b_{1} \
&\vdots \
g_{M}(x) &= b_{M}
\end{align}
$$
The key condition we’re examining is the gradient linear combination:
$$
\nabla f(x) = \sum_{i=1}^{M} \lambda_{i}\nabla g_{i}(x)
$$
The Core Reason: Geometry + Local Optimality
Let’s break this down into three concrete steps:
1. The Feasible Set’s Tangent Space
First, define the feasible set $S = {x \in \mathbb{R}^n \mid g_i(x) = b_i \text{ for all } i=1..M}$. At any feasible point $x^$, the tangent space $T_{x^}S$ is the set of all vectors $v$ that represent infinitesimal moves you can make from $x^*$ while staying on $S$.
For a vector $v$ to be in this tangent space, it must satisfy $\nabla g_i(x^) \cdot v = 0$ for every $i$. Why? Because the gradient of $g_i$ is always perpendicular to its level set $g_i(x)=b_i$—so any direction tangent to the level set can’t have a component aligned with $\nabla g_i(x^)$.
2. Local Maximum = No Feasible Improvements
If $x^$ is a local maximum of $f$ on $S$, there can be no feasible direction $v$ (from $T_{x^}S$) where moving along $v$ increases $f$. Mathematically, this means the directional derivative of $f$ along every $v \in T_{x^}S$ must be $\leq 0$:
$$
\nabla f(x^) \cdot v \leq 0 \quad \forall v \in T_{x^}S
$$
But here’s a critical point: if $v$ is a feasible direction, so is $-v$ (moving the opposite way along the tangent space). For the inequality to hold for both $v$ and $-v$, the directional derivative must be zero for all $v \in T_{x^}S$:
$$
\nabla f(x^) \cdot v = 0 \quad \forall v \in T_{x^}S
$$
In plain terms: $\nabla f(x^)$ must be perpendicular to every possible feasible move from $x^$.
3. Orthogonal Complement = Span of Constraint Gradients
From linear algebra, the set of all vectors perpendicular to $T_{x^}S$ is called the orthogonal complement of $T_{x^}S$. Now, if we assume the linear independence constraint qualification (LICQ)—meaning $\nabla g_1(x^), ..., \nabla g_M(x^)$ are linearly independent—then this orthogonal complement is exactly the span of the constraint gradients.
What does that mean? It means $\nabla f(x^)$ has to lie within the set of all possible linear combinations of $\nabla g_1(x^), ..., \nabla g_M(x^)$. By definition of a span, there must exist scalars $\lambda_1, ..., \lambda_M$ (our Lagrange multipliers) such that:
$$
\nabla f(x^) = \sum_{i=1}^M \lambda_i \nabla g_i(x^*)
$$
Quick Intuition Recap
To simplify: At a local maximum, the steepest ascent direction of $f$ (given by $\nabla f$) can’t point into the feasible set—because if it did, you could move that way to get a higher $f$ value. Instead, $\nabla f$ has to point directly perpendicular to all feasible moves. Since each constraint’s gradient is perpendicular to its own level set, the only way $\nabla f$ is perpendicular to every feasible direction is if it’s a mix (linear combination) of those constraint gradients.
A Key Caveat
This all relies on LICQ: if the constraint gradients are linearly dependent at $x^$, the orthogonal complement might be larger than the span of the gradients, and the linear combination condition might not hold in this exact form. But LICQ is a standard assumption in most rigorous treatments of Lagrange multipliers because it ensures the feasible set behaves "nicely" at $x^$.
内容的提问来源于stack exchange,提问作者sergio.azevedo

