You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

拉格朗日乘数法:为何目标函数梯度是约束梯度的线性组合?

Why Does $\nabla f(x)$ Equal a Linear Combination of Constraint Gradients in Lagrange Multipliers?

Great question—let’s unpack this rigorously, starting from the geometry of constrained optimization and working our way to the formal calculus condition.

First, Let’s Restate the Problem

We’re targeting a local maximum of $f(x)$ under equality constraints, formalized as:
$$
\begin{align}
\max:, f(x),& \enspace x \in \mathbb{R}^{n} \
\text{s.t.} \quad g_{1}(x) &= b_{1} \
&\vdots \
g_{M}(x) &= b_{M}
\end{align}
$$
The key condition we’re examining is the gradient linear combination:
$$
\nabla f(x) = \sum_{i=1}^{M} \lambda_{i}\nabla g_{i}(x)
$$

The Core Reason: Geometry + Local Optimality

Let’s break this down into three concrete steps:

1. The Feasible Set’s Tangent Space

First, define the feasible set $S = {x \in \mathbb{R}^n \mid g_i(x) = b_i \text{ for all } i=1..M}$. At any feasible point $x^$, the tangent space $T_{x^}S$ is the set of all vectors $v$ that represent infinitesimal moves you can make from $x^*$ while staying on $S$.

For a vector $v$ to be in this tangent space, it must satisfy $\nabla g_i(x^) \cdot v = 0$ for every $i$. Why? Because the gradient of $g_i$ is always perpendicular to its level set $g_i(x)=b_i$—so any direction tangent to the level set can’t have a component aligned with $\nabla g_i(x^)$.

2. Local Maximum = No Feasible Improvements

If $x^$ is a local maximum of $f$ on $S$, there can be no feasible direction $v$ (from $T_{x^}S$) where moving along $v$ increases $f$. Mathematically, this means the directional derivative of $f$ along every $v \in T_{x^}S$ must be $\leq 0$:
$$
\nabla f(x^
) \cdot v \leq 0 \quad \forall v \in T_{x^}S
$$
But here’s a critical point: if $v$ is a feasible direction, so is $-v$ (moving the opposite way along the tangent space). For the inequality to hold for both $v$ and $-v$, the directional derivative must be zero for all $v \in T_{x^
}S$:
$$
\nabla f(x^) \cdot v = 0 \quad \forall v \in T_{x^}S
$$
In plain terms: $\nabla f(x^)$ must be perpendicular to every possible feasible move from $x^$.

3. Orthogonal Complement = Span of Constraint Gradients

From linear algebra, the set of all vectors perpendicular to $T_{x^}S$ is called the orthogonal complement of $T_{x^}S$. Now, if we assume the linear independence constraint qualification (LICQ)—meaning $\nabla g_1(x^), ..., \nabla g_M(x^)$ are linearly independent—then this orthogonal complement is exactly the span of the constraint gradients.

What does that mean? It means $\nabla f(x^)$ has to lie within the set of all possible linear combinations of $\nabla g_1(x^), ..., \nabla g_M(x^)$. By definition of a span, there must exist scalars $\lambda_1, ..., \lambda_M$ (our Lagrange multipliers) such that:
$$
\nabla f(x^
) = \sum_{i=1}^M \lambda_i \nabla g_i(x^*)
$$

Quick Intuition Recap

To simplify: At a local maximum, the steepest ascent direction of $f$ (given by $\nabla f$) can’t point into the feasible set—because if it did, you could move that way to get a higher $f$ value. Instead, $\nabla f$ has to point directly perpendicular to all feasible moves. Since each constraint’s gradient is perpendicular to its own level set, the only way $\nabla f$ is perpendicular to every feasible direction is if it’s a mix (linear combination) of those constraint gradients.

A Key Caveat

This all relies on LICQ: if the constraint gradients are linearly dependent at $x^$, the orthogonal complement might be larger than the span of the gradients, and the linear combination condition might not hold in this exact form. But LICQ is a standard assumption in most rigorous treatments of Lagrange multipliers because it ensures the feasible set behaves "nicely" at $x^$.

内容的提问来源于stack exchange,提问作者sergio.azevedo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:05:37