如何从概念层面理解分词词与词嵌入的关系及映射机制
Hey there! Let me break this down for you step by step—this is a super common point of confusion when working with word embeddings and text models, so you’re not alone.
Understanding the Link Between Tokenized Sequences and Word Embeddings
1. First, Let’s Clarify the Two Core Pieces
Let’s start by separating the two things you’re confused about:
- Tokenized sequences (
x_train): These are just lists of integers. Each integer maps to a specific word in a predefined vocabulary. For example, the sentence "The cat sat" might get tokenized to[5, 12, 27], where 5 = "The", 12 = "cat", 27 = "sat". - Word embeddings: This is a matrix where each row represents a word as a dense numerical vector. The row index in this matrix exactly matches the integer token from your sequence. So row 12 in the embedding matrix is the vector representation for "cat".
2. The Vocabulary: The Glue Between Tokens and Embeddings
The key to this mapping is the vocabulary dictionary—a lookup table created before you tokenize your data. Here’s how it works when you use the entire dataset to build it:
- First, you scan every piece of text (train, validation, test) to count word frequencies.
- You then select the top N most frequent words (e.g., 10,000) to form your vocabulary. Each word gets assigned a unique integer index (starting from 0 or 1).
- This vocabulary is fixed once created. When you tokenize
x_train, every word in the training data gets replaced by its corresponding index from this dictionary. Words not in the vocabulary get mapped to an "unknown" token (usually index 0 or vocab_size).
3. The Mapping Mechanism: From Token Index to Embedding Vector
Once you have your tokenized sequences and vocabulary, the embedding layer in your model handles the actual mapping:
- The embedding layer is initialized as a matrix with shape
(vocab_size, embedding_dim)(e.g., 10,000 rows × 128 columns for 128-dimensional embeddings). - When you feed
x_train(integer sequences) into the embedding layer, it performs a simple lookup: for each integer index in the sequence, it pulls the corresponding row from the embedding matrix. - For example, if your token sequence is
[5, 12, 27], the embedding layer outputs a 3×128 tensor where each row is the embedding vector for the word at index 5, 12, and 27 respectively.
4. Why Using the Entire Dataset for Vocabulary Is Safe (and Smart)
You might worry that using test data to build the vocabulary is "data leakage"—but it’s not, and here’s why:
- The vocabulary only defines a fixed mapping from words to integers. It doesn’t contain any information about the labels or the task you’re training on.
- While you’re using test data to include more words in your vocabulary, the embedding layer’s weights (the actual vector values) are only updated using the training data. Words that appear only in the test set will have their embedding vectors stay at their initial random values (or pre-trained values if you’re using pre-trained embeddings) during training.
- This approach reduces the number of "unknown" tokens in your test set, which helps your model generalize better to words it didn’t see during training.
5. A Quick Concrete Example
Let’s make this tangible:
- Vocabulary creation: You scan all data and create a vocab:
{"the": 0, "cat": 1, "sat": 2, "dog": 3} - Tokenization: Your training sentence "The cat sat" becomes
[0, 1, 2] - Embedding matrix initialization:
[[0.1, 0.3], # "the" (index 0) [0.2, 0.4], # "cat" (index 1) [0.5, 0.7], # "sat" (index 2) [0.6, 0.8]] # "dog" (index 3, only in test data) - Mapping in the model: When you feed
[0,1,2]into the embedding layer, it outputs:[[0.1, 0.3], [0.2, 0.4], [0.5, 0.7]] - Training: During training, only the first 3 rows of the embedding matrix (corresponding to training words) get updated. The 4th row ("dog") stays unchanged until test time.
内容的提问来源于stack exchange,提问作者sparkh2o
相关产品推荐
相关产品推荐

