孪生网络(Siamese networks)为何需复制网络?基于DeepFace的技术疑问
Great question—this is a super common point of confusion when first diving into siamese networks, so let’s break this down clearly:
1. It’s About Computational Graph & Batch Efficiency
In modern deep learning frameworks (like TensorFlow or PyTorch), we need to process two input images simultaneously (e.g., an anchor face and a positive/negative match) to calculate metric losses like the contrastive loss used in DeepFace.
If you tried to use a single DNN sequentially—processing image A, saving its features, then processing image B—you’d force the framework to track two separate forward passes when computing backpropagation. This is inefficient, especially for batch training, and complicates the automatic differentiation pipeline.
The "copied" networks you see are actually two branches of the same weight graph—not duplicate parameter tensors in memory. Frameworks handle this by reusing the same set of weights across both branches, letting you process both inputs in a single forward pass. This builds a clean, unified computational graph that’s easy for the framework to optimize and backpropagate through.
2. Clarity in Loss Definition & Architecture
Metric losses depend directly on comparing the feature embeddings of two inputs. Using two shared-weight branches makes the loss calculation explicit: loss = compute_loss(f(x1), f(x2), label), where f is the exact same feature extractor for both x1 and x2.
If you used a single DNN sequentially, you’d have to manually cache the first embedding, compute the second, then calculate the loss. This adds unnecessary complexity, and the "copied" structure is far more intuitive for communicating how the network works (both in papers and code). It’s a conceptual shorthand to show that the same feature extraction logic applies to both inputs.
3. Historical & Conceptual Precedent
Early siamese network papers (like the 2005 one from LeCun et al.) visualized the architecture with two parallel, weight-sharing branches. This became the standard way to represent these networks, even though the underlying implementation doesn’t duplicate parameters. DeepFace followed this convention for consistency with existing literature—readers familiar with siamese networks would immediately recognize the structure.
To Answer Your Final Question:
Yes, they’re essentially referring to the same thing as using one DNN for both inputs! The "copy" wording is just a structural description, not a literal duplication of parameters. In practice, you’d implement this by reusing the same model instance for both inputs (e.g., in PyTorch, passing both images through the same nn.Module and computing the loss on their outputs) or using framework tools to explicitly share weights across two branches.
内容的提问来源于stack exchange,提问作者Simon Hessner

