You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

技术问询:深度学习‘Norm’术语含义及正则化场景关联疑问

Hey there! Let's break this down clearly for you.

What does the term 'Norm' mean in neural networks?

In the context of neural networks, a norm is a mathematical measure of the "size" or "length" of a vector (like a hidden state, weight vector, or activation vector). The most common types you'll encounter are:

  • L1 Norm: Calculated as the sum of the absolute values of the vector's elements. For a vector [x1, x2, ..., xn], it's |x1| + |x2| + ... + |xn|. It's often used for sparsity-inducing regularization.
  • L2 Norm: Calculated as the square root of the sum of the squared elements of the vector. For the same vector, it's sqrt(x1² + x2² + ... + xn²). This is the classic Euclidean distance you might remember from geometry, and it's widely used to measure vector magnitude in neural networks.

There are other norms too (like L0, L∞), but L1 and L2 are the most prevalent in practice.

What does 'Norm' refer to in that RNN regularization paper?

First, let's look at the original sentence you shared:

We stabilize the activations of Recurrent Neural Networks (RNNs) by penalizing the squared distance between successive hidden states’ norms

In this context, the "norms" being referenced are almost certainly the L2 Norms of the RNN's hidden state vectors. Here's why:

  • RNN hidden states are high-dimensional vectors, and their L2 norm gives a straightforward measure of their overall magnitude. Penalizing the squared difference between consecutive hidden state norms means the model is encouraged to keep the size of its hidden states consistent across time steps. This helps stabilize training by preventing hidden states from growing or shrinking drastically (a common cause of vanishing/exploding gradients in RNNs).
  • The mention of "squared distance" aligns perfectly with L2 norms—since the squared L2 norm avoids computing a square root, it's computationally cheaper and often used in loss/penalty terms for efficiency.

Now, let's connect this to the other techniques you asked about:

  • L1/L2 Normalization: These are operations where you divide a vector by its L1 or L2 norm to scale it to a fixed length (e.g., L2 normalization makes the vector have an L2 norm of 1). The paper's use of "norm" is about measuring the vector's magnitude for regularization, not performing a normalization operation—but it's still rooted in the same L2 norm math.
  • Batch Normalization: This is a completely different technique. Batch Norm normalizes the distribution of activations across a batch of data (subtracting the batch mean and dividing by the batch standard deviation) to speed up training. It has no direct connection to the vector norm being discussed here; it's about stabilizing feature distributions, not vector magnitudes.

内容的提问来源于stack exchange,提问作者Kari

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:33:12