You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

斯坦福NLP课程SkipGram中上下文词表示矩阵含义及相关概念问询

Understanding SkipGram's Context Word Representation Matrix

Hey there! Let's unpack your confusion about the SkipGram components mentioned in Stanford's NLP course—those matrix details can feel tricky at first, but we'll break them down clearly.

Let's Walk Through Each Component You Mentioned

  • One-hot encoder vector (first column)
    This is your input vector for the center word, with dimensions V×1 where V is the size of your vocabulary. It's a sparse vector: only the index corresponding to your center word is set to 1, and every other entry is 0. For example, if your vocabulary has 5000 words and your center word is "dog", the 123rd index (say that's where "dog" lives in your vocab) is 1, and all 4999 other positions are 0.

  • Word embedding matrix (second column)
    This is a dense matrix W with dimensions V×d (where d is your embedding dimension—think 100, 300, etc.). Each row in this matrix represents the embedding vector for a single word in your vocabulary. When you multiply the one-hot vector by this matrix, you're essentially "picking out" the row corresponding to your center word—since the one-hot vector only has a 1 in that position, the result is exactly that word's embedding.

  • Word representation vector (third column)
    This is the output of multiplying the one-hot input by the embedding matrix W, so it's a d×1 vector. Put simply, this is the dense, low-dimensional representation of your center word that the model learns. You might see this written as h = W^T * x where x is the one-hot vector—this just transposes the matrix to get a column vector instead of a row.

  • V×d context word representation matrix (fourth column)
    Now this is where your confusion probably hits. Let's clarify:

    • In the original SkipGram setup, this is a separate matrix W' (also V×d). Each row in W' represents a context word's "score vector". The model uses the center word's embedding h to calculate a score for every possible word in the vocab (to predict which ones are context words), and that score comes from multiplying h with each row in W'. Then a softmax layer turns those scores into probabilities of each word being a context word for the center word.
    • The part about "three matrices being identical" likely refers to weight sharing—a common optimization or simplification where we set W = W'. In this case, the same matrix is used to generate both center word embeddings and context word score vectors. So the center word's embedding is a row from W, and each context word's score vector is also a row from the same W. This reduces the number of parameters the model needs to learn, which can speed up training and prevent overfitting on smaller datasets.

Quick Recap to Tie It All Together

The one-hot vector acts as a pointer to grab the center word's embedding from the W matrix. If we share weights, that same W matrix is also used to compute scores for potential context words. Without sharing, W and W' are separate matrices, but the core idea remains: transforming sparse one-hot inputs into dense embeddings, then using those embeddings to predict context words.

内容的提问来源于stack exchange,提问作者Bratt Swan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:04:21