请求详解w2v(Word2Vec CBOW算法)中的矩阵与向量维度
Great question—let’s break down the matrix and vector dimensions in Word2Vec’s CBOW model step by step, since this is a common point of confusion when first diving into word embeddings. I’ll keep it practical, with concrete examples and why each choice matters.
1. Core Matrices & Vectors: Definitions and Dimensions
First, let’s recap CBOW’s basic logic: it predicts a center word using its surrounding context words. Here are the key players and their dimensions:
Input Embedding Matrix (
W)
Dimension:V × DV= size of your vocabulary (number of unique words in your corpus, after filtering low-frequency terms)D= word vector dimension (a hyperparameter you choose, e.g., 100, 200, 300)
What it is: Each row in this matrix is the input embedding vector for a single word. When you represent a word as a one-hot vector (sizeV), multiplying it byWpulls out the corresponding row—your D-dimensional word embedding. For example, if you have 10,000 unique words and setD=200,Wis a 10,000×200 matrix.
Context Vector (
c)
Dimension:D × 1(column vector)
What it is: CBOW takes all the input embeddings of the context words, averages them (or sums them—implementation varies slightly), and outputs this single D-dimensional vector. It’s the condensed "semantic fingerprint" of the context, used to guess the center word.Output Projection Matrix (
W')
Dimension:D × V
What it is: This matrix maps the D-dimensional context vector into a space the size of your vocabulary. Multiplying the context vectorcbyW'gives you aV × 1vector where each element is a "score" for how likely that word is the center word. We then run this through softmax to turn scores into probabilities.Target One-Hot Vector
Dimension:V × 1
What it is: The ground-truth representation of the center word—only the position corresponding to the center word is 1, everything else is 0. We use this to calculate cross-entropy loss during training.
2. Why These Dimensions?
None of these choices are arbitrary—they’re driven by both logic and practicality:
Vocabulary size
V: This is a hard constraint from your data. You can’t represent a unique word with a one-hot vector smaller thanV, so all matrices tied to word identification (input matrix rows, output matrix columns) must matchV. Most implementations let you filter low-frequency words to shrinkV(since rare words don’t get enough training data anyway).Word vector dimension
D: This is a balancing act:- Too small (
D=50for a large corpus): Your embeddings won’t capture nuanced semantic relationships—"king" and "queen" might end up with nearly identical vectors, or "run" (verb) and "run" (noun) won’t be distinguishable. - Too large (
D=1000for a small corpus): You’ll waste compute resources (10k words × 1000 dimensions = 10 million parameters!) and risk overfitting—your model will memorize rare word pairs instead of learning generalizable patterns.
Industry defaults are 100–300; adjust based on corpus size (bigger corpus = biggerD).
- Too small (
Context vector’s
Ddimension: It has to match the input embedding dimension to act as a bridge between the input space and output space. Averaging D-dimensional vectors keeps the dimension consistent, so it can multiply cleanly with theD×Voutput matrix.
3. Critical Notes & Pitfalls
WandW'are not transposes!
A common mistake is thinking these two matrices are inverses or transposes—they’re separate parameters trained independently. In practice, most tools (like gensim’s Word2Vec) only keep the input matrixWas the final word embeddings, since they do a better job capturing semantic similarity.Optimizations like negative sampling don’t change core dimensions
The original CBOW uses full softmax, which is slow for largeV(you have to compute scores for all words). Negative sampling or hierarchical softmax speed this up by avoiding fullVcalculations, but the input matrixV×Dstays the same—you’re just changing how you compute the loss, not the core embedding structure.Dimension alignment is non-negotiable
If your one-hot vector is1×V, your input matrix must beV×Dto get a1×Dembedding. If your context vector isD×1, your output matrix must beD×Vto get aV×1score vector. Mess this up, and your code will throw a matrix multiplication error immediately.Low-frequency words waste dimension space
Even if you setD=300, a word that only appears 3 times in your corpus won’t have enough data to learn a meaningful embedding. Always filter low-frequency terms first—this shrinksV, reduces compute, and makes your high-quality embeddings more reliable.
内容的提问来源于stack exchange,提问作者MrCorwin160

