斯坦福NLP课程SkipGram中上下文词表示矩阵含义及相关概念问询
Hey there! Let's unpack your confusion about the SkipGram components mentioned in Stanford's NLP course—those matrix details can feel tricky at first, but we'll break them down clearly.
Let's Walk Through Each Component You Mentioned
One-hot encoder vector (first column)
This is your input vector for the center word, with dimensionsV×1whereVis the size of your vocabulary. It's a sparse vector: only the index corresponding to your center word is set to1, and every other entry is0. For example, if your vocabulary has 5000 words and your center word is "dog", the 123rd index (say that's where "dog" lives in your vocab) is1, and all 4999 other positions are0.Word embedding matrix (second column)
This is a dense matrixWwith dimensionsV×d(wheredis your embedding dimension—think 100, 300, etc.). Each row in this matrix represents the embedding vector for a single word in your vocabulary. When you multiply the one-hot vector by this matrix, you're essentially "picking out" the row corresponding to your center word—since the one-hot vector only has a1in that position, the result is exactly that word's embedding.Word representation vector (third column)
This is the output of multiplying the one-hot input by the embedding matrixW, so it's ad×1vector. Put simply, this is the dense, low-dimensional representation of your center word that the model learns. You might see this written ash = W^T * xwherexis the one-hot vector—this just transposes the matrix to get a column vector instead of a row.V×d context word representation matrix (fourth column)
Now this is where your confusion probably hits. Let's clarify:- In the original SkipGram setup, this is a separate matrix
W'(alsoV×d). Each row inW'represents a context word's "score vector". The model uses the center word's embeddinghto calculate a score for every possible word in the vocab (to predict which ones are context words), and that score comes from multiplyinghwith each row inW'. Then a softmax layer turns those scores into probabilities of each word being a context word for the center word. - The part about "three matrices being identical" likely refers to weight sharing—a common optimization or simplification where we set
W = W'. In this case, the same matrix is used to generate both center word embeddings and context word score vectors. So the center word's embedding is a row fromW, and each context word's score vector is also a row from the sameW. This reduces the number of parameters the model needs to learn, which can speed up training and prevent overfitting on smaller datasets.
- In the original SkipGram setup, this is a separate matrix
Quick Recap to Tie It All Together
The one-hot vector acts as a pointer to grab the center word's embedding from the W matrix. If we share weights, that same W matrix is also used to compute scores for potential context words. Without sharing, W and W' are separate matrices, but the core idea remains: transforming sparse one-hot inputs into dense embeddings, then using those embeddings to predict context words.
内容的提问来源于stack exchange,提问作者Bratt Swan

