You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Mikolov 2010年RNN语言模型中词嵌入$w(t)$的获取方式问询

Understanding Word Embedding $w(t)$ in Mikolov 2010's RNNLM

Great question—let’s unpack this clearly, since the connection between Bengio’s 2003 A Neural Probabilistic Language Model and Mikolov’s 2010 Recurrent neural network based language model is key to understanding how word embeddings evolved here.

First, let’s recap what you already noted about Bengio’s 2003 model to set the stage:

In Bengio et al.'s 2003 paper, each word gets its own independent embedding vector, stored in a trainable matrix. This "semi-layer" is just a lookup table—no nonlinear transformation happens here; it maps word indices directly to their vectors. During training, this embedding matrix is updated alongside other model parameters via backpropagation, as the model learns to minimize language modeling loss.

Now, onto Mikolov’s 2010 RNNLM: the core approach to word embeddings $w(t)$ builds on Bengio’s work but adapts it for recurrent architectures:

  • Initialization: Just like Bengio’s model, $w(t)$ starts as a vector pulled from a trainable embedding matrix. Each word in the vocabulary gets an initial random (or optionally pre-initialized) vector, and the embedding layer still acts as a simple lookup with no built-in nonlinearity.
  • Integration with RNN: The big difference is how this embedding is used. In the RNNLM, $w(t)$ (the embedding of the current word at time step $t$) is concatenated with the previous hidden state $h(t-1)$. This combined vector feeds into the recurrent layer to compute the new hidden state $h(t)$ (usually using a tanh activation).
  • Training & Updates: Critically, the embedding matrix is still updated during backpropagation—just like in Bengio’s model. The cross-entropy loss from predicting the next word flows back through the RNN layers, the hidden state calculations, and finally into the embedding matrix. This means $w(t)$ isn’t static; it’s learned jointly with the RNN’s weights to optimize the language modeling task.

One quick side note: Mikolov’s later work (like Word2Vec) would optimize embeddings more directly, but the 2010 RNNLM still ties embedding learning tightly to the language modeling objective—similar to Bengio’s approach, just within a recurrent framework that captures sequential context more effectively.

To sum it up: $w(t)$ in Mikolov 2010 is a trainable word embedding, initialized as a lookup vector and updated via backpropagation, just like in Bengio’s 2003 model. The main distinction is how it’s fed into the recurrent layer instead of a feedforward network.

内容的提问来源于stack exchange,提问作者Ziemo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:54:09