关于torch.Tensor.add_方法的疑问及PyTorch SGD源码困惑
Tensor.add_ Overloads and SGD's Usage Great question—this trips up a lot of folks because PyTorch's tensor methods have multiple overloaded signatures that aren't always obvious at first glance. Let's break this down step by step.
Why the SGD code works, but your test doesn't
First, let's clarify what's happening in the SGD source code line:
p.data.add_(-group['lr'], d_p)
This isn't passing two tensors—it's passing a scalar (-group['lr']) and a tensor (d_p). PyTorch has a specific overload of add_ for this exact scenario:
add_(value: Number, other: Tensor) → TensorPerforms the operation:
self = self + value * otherin-place.
Which is exactly what SGD needs: param = param - lr * grad (since -group['lr'] is the scalar multiplier for the gradient tensor d_p).
Your test code, on the other hand, tries to pass two tensors to add_:
a.add_(b, c) # Error! No overload accepts two tensors as positional args
There's no add_ signature that takes two positional tensors—hence the parameter count error.
The different add_ signatures you need to know
PyTorch's add_ has a few key overloads to handle different use cases:
- Scalar + Tensor (positional):
This is the one used in SGD.tensor.add_(scalar, other_tensor) # Equivalent to: tensor += scalar * other_tensor - Tensor + Alpha (keyword-only):
This is the modern, more readable way to do the same operation as above (using Python's keyword-only arguments to avoid ambiguity).tensor.add_(other_tensor, alpha=scalar) # Equivalent to: tensor += scalar * other_tensor - Simple in-place addition:
This is the most straightforward overload—just add another tensor directly.tensor.add_(other_tensor) # Equivalent to: tensor += other_tensor
Fixing your test code
If you want to replicate a scalar-multiplied addition like the SGD code, use either the scalar-first positional form (if the scalar is a number) or the keyword alpha argument. For your example, if you want a = a + b * c, you'd write:
import torch a = torch.tensor([1,2,3]) b = torch.tensor([6,10,15]) c = torch.tensor([100,100,100]) a.add_(c, alpha=b) # Uses the keyword-only alpha overload print(a) # Output: tensor([601, 1002, 1503])
How add_ works under the hood
As an in-place operation (marked by the _ suffix), add_ directly modifies the underlying memory of the tensor it's called on, rather than creating a new tensor. This is crucial for optimizers like SGD because it saves memory—we don't need to allocate new tensors for every parameter update.
PyTorch dispatches to the correct overload based on the types and number of arguments you pass. The scalar-first overload exists specifically for common optimization operations (like applying a learning rate multiplier to a gradient), which is why it's used in the SGD implementation.
内容的提问来源于stack exchange,提问作者Daniel Möller

