如何计算自适应梯度(AdaGrad)?FaceNet中Triplet Loss梯度计算疑问
Hey there! Let me break down AdaGrad and how it applies to computing gradients for Triplet Loss—since that's what you're stuck on after diving into FaceNet. I'll start with the basics of AdaGrad, then tie it directly to the Triplet Loss use case in FaceNet.
First, What is AdaGrad?
AdaGrad short for Adaptive Gradient is a gradient descent optimization algorithm that automatically adjusts the learning rate for each individual model parameter. The core idea is simple: parameters that get updated frequently (i.e., have large gradients) get a smaller learning rate over time, while parameters that are updated rarely (sparse gradients) keep a larger learning rate. This is super useful for tasks like face recognition, where some features are more critical than others.
Here's the formal update rule in plain terms:
- Calculate the gradient ( g_t ) of the loss with respect to a parameter ( w_t ) at step ( t ).
- Maintain a running sum ( G_t ) of the squares of all past gradients for that parameter: ( G_t = G_{t-1} + g_t^2 ).
- Update the parameter using a scaled learning rate:
[
w_{t+1} = w_t - \frac{\eta}{\sqrt{G_t + \epsilon}} \times g_t
]- ( \eta ) is your initial, global learning rate.
- ( \epsilon ) is a tiny value (like 1e-8) to avoid dividing by zero when ( G_t ) is small.
The key here is that ( \sqrt{G_t} ) grows as we accumulate more gradient squares, so the effective learning rate shrinks for parameters that have been updated a lot.
How AdaGrad Works with Triplet Loss
First, let's recap Triplet Loss since it's the loss function FaceNet uses:
[
L = \max(0, d(a,p) - d(a,n) + \alpha)
]
Where:
- ( d(a,p) ) = distance between an anchor face ( a ) and a positive face ( p ) same person
- ( d(a,n) ) = distance between the anchor ( a ) and a negative face ( n ) different person
- ( \alpha ) = a margin that ensures the positive distance is always smaller than the negative distance by at least this amount
Step 1: Compute Triplet Loss Gradients
AdaGrad only kicks in after we calculate the gradient of the Triplet Loss with respect to the model's parameters (like convolution weights, fully connected layer biases, etc.). Here's when gradients matter:
- If ( d(a,p) - d(a,n) + \alpha \leq 0 ): The loss is 0, so there's no gradient—parameters don't update.
- If ( d(a,p) - d(a,n) + \alpha > 0 ): The loss is positive, so the gradient is ( \nabla d(a,p) - \nabla d(a,n) ) (since the max function's gradient is 1 when its input is positive).
Step 2: Apply AdaGrad to Update Parameters
Once we have that gradient ( g_t ) for each parameter, we feed it into the AdaGrad update rule:
- For every parameter in the model, we add the square of its current gradient to its running sum ( G_t ).
- We scale the initial learning rate ( \eta ) by ( 1/\sqrt{G_t + \epsilon} ), then multiply by the gradient to get the parameter update.
- We apply this scaled update to the parameter.
Why This Makes Sense for FaceNet
FaceNet's goal is to learn compact, discriminative face embeddings. AdaGrad is a great fit here because:
- Sparse gradients: Some model parameters will be responsible for detecting critical facial features (like eyes, nose) and will have large, frequent gradients. AdaGrad slows down their updates to prevent overshooting.
- Fine-grained features: Other parameters might capture subtle features (like skin texture) that only show up in rare cases. AdaGrad keeps their learning rate high, so the model can still learn these nuances.
- No manual tuning: You don't have to manually set learning rates for each parameter—AdaGrad handles the adaptation automatically, which is a huge plus for large-scale models like FaceNet.
Simplified Pseudocode to Illustrate
Here's a quick pseudocode snippet to make this concrete using PyTorch-like syntax:
# Initialize model parameters and AdaGrad state model = FaceNetModel() params = list(model.parameters()) # G stores the running sum of squared gradients for each parameter gradient_sums = {param: torch.zeros_like(param) for param in params} initial_lr = 0.01 epsilon = 1e-8 alpha = 0.2 # Triplet Loss margin for anchor, positive, negative in training_data: # Compute embeddings and distances emb_a = model(anchor) emb_p = model(positive) emb_n = model(negative) d_ap = torch.norm(emb_a - emb_p, p=2) d_an = torch.norm(emb_a - emb_n, p=2) # Calculate Triplet Loss loss = torch.max(torch.tensor(0.0), d_ap - d_an + alpha) if loss > 0: # Compute gradients of loss w.r.t. model parameters loss.backward() # Apply AdaGrad updates for param in params: grad = param.grad # Update running sum of squared gradients gradient_sums[param] += grad ** 2 # Compute scaled learning rate scaled_lr = initial_lr / torch.sqrt(gradient_sums[param] + epsilon) # Update parameter param.data -= scaled_lr * grad # Reset gradient for next iteration param.grad.zero_()
A Quick Note on AdaGrad's Limitation
One thing to keep in mind: since ( G_t ) is a cumulative sum, it keeps growing indefinitely. This means the effective learning rate can eventually shrink to near zero, causing training to stall. Later algorithms like RMSProp and Adam fix this by using a moving average instead of a cumulative sum, but FaceNet uses the original AdaGrad—so just be mindful of your initial learning rate choice!
内容的提问来源于stack exchange,提问作者mirantha

