You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于Annotated Transformer Encoder LayerNorm实现正确性的问询

《Annotated Transformer》的Encoder LayerNorm实现与原论文对比
  • Transformer原论文规定的子层输出逻辑为:LayerNorm(x + Dropout(SubLayer(x))),即LayerNorm要在**Dropout(子层输出)**之后应用,属于「后归一化」结构。
  • 但《Annotated Transformer》中的实现将LayerNorm置于子层之前,逻辑为:x + Dropout(SubLayer(LayerNorm(x))),与原论文顺序完全相反,对应代码如下:
class SublayerConnection(nn.Module):
    """
    A residual connection followed by a layer norm.
    Note for code simplicity the norm is first as opposed to last.
    """

    def __init__(self, size, dropout):
        super(SublayerConnection, self).__init__()
        self.norm = LayerNorm(size)
        self.dropout = nn.Dropout(dropout)

    def forward(self, x, sublayer):
        "Apply residual connection to any sublayer with the same size."
        return x + self.dropout(sublayer(self.norm(x)))   # <--- LayerNorm before SubLayer

从严格贴合原论文设计的角度看,这个实现不符合规定,但代码注释已明确说明「为了代码简洁将归一化放在前面而非后面」,属于故意采用的变体。这种「预归一化」的方式后续在实践中被证明训练更稳定,如今不少Transformer衍生架构也会采用,但确实和原论文的原始设计存在差异。

内容的提问来源于stack exchange,提问作者mon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 12:02:43