如何基于训练好的Gensim Word2Vec模型添加新词并保留旧词向量?
Great question! I’ve run into this exact scenario before—wanting to extend a pre-trained Word2Vec model with new words without messing up the existing embeddings. Here are a few practical approaches that align with your request for a doc2vec-infer-like solution:
1. Infer New Word Vectors Using Pre-Trained Context (Most Similar to Doc2Vec Infer)
This method mimics doc2vec's inference process: we fix the existing word vectors and optimize a new vector for each unseen word until it fits naturally with its context from your new dataset.
How it works:
- For each new word in your dataset, collect all its context windows (match the window size used to train your original Word2Vec model).
- Initialize a random vector for the new word.
- Use the original model's training objective (CBOW or Skip-gram) to tweak this vector:
- For CBOW: Adjust the new word's vector so that it predicts the average of its context vectors.
- For Skip-gram: Adjust the new word's vector so that its context words are predicted from it.
- Stop when the vector stabilizes (or after a fixed number of epochs, like doc2vec's
infer_epochs).
Example Code Snippet:
import numpy as np from gensim.models import Word2Vec # Load your pre-trained model original_model = Word2Vec.load("path/to/your/model") wv = original_model.wv vector_size = wv.vector_size window = original_model.window learning_rate = 0.01 # Adjust based on your needs def infer_new_word_vector(new_word, context_windows): # Initialize random vector new_vec = np.random.randn(vector_size) * 0.01 for _ in range(10): # Mimic infer_epochs total_loss = 0.0 for context in context_windows: # Filter out words not in original model (we only trust pre-trained vectors) valid_context = [word for word in context if word in wv] if not valid_context: continue # CBOW objective: new_vec should predict the average context vector target_vec = np.mean(wv[valid_context], axis=0) # Calculate error error = target_vec - new_vec # Update new_vec new_vec += learning_rate * error total_loss += np.sum(error**2) if total_loss < 1e-4: # Stop early if loss is small break return new_vec # Usage: Collect context windows for your new word new_word = "newword" context_windows = [["oldword1", "oldword2"], ["oldword3", "newword", "oldword4"]] # Example contexts new_vec = infer_new_word_vector(new_word, context_windows) # Add the new word to the model's KeyedVectors wv.add_vectors([new_word], [new_vec]) # Save the updated model original_model.save("path/to/updated/model")
Notes:
- This preserves all original embeddings since we never modify them.
- The quality depends heavily on how many valid context windows you have for each new word (more contexts = better vectors).
- You can adapt this for Skip-gram by reversing the objective (predict context words from the new word vector).
2. Incremental Training with Frozen Original Embeddings
Gensim's default update=True mode will adjust all vectors (old and new), but we can hack the model to freeze original words while training new ones.
How it works:
- Load your pre-trained model and extend its vocabulary with the new dataset.
- Freeze the vectors of existing words by marking them as "untrainable" (we'll modify the model's internal weights to prevent updates).
- Train only on the new dataset, so only new word vectors are updated.
Example Code Snippet:
from gensim.models import Word2Vec # Load pre-trained model original_model = Word2Vec.load("path/to/your/model") old_vocab = set(original_model.wv.key_to_index.keys()) # Extend vocabulary with new corpus new_corpus = [["newword", "oldword1"], ["oldword2", "newword"]] # Your new dataset original_model.build_vocab(new_corpus, update=True) # Freeze original word vectors: save old vectors, then restore them after each training batch old_vectors = original_model.wv.vectors.copy() # Override the train method to restore old vectors after each update (simplified approach) def custom_train(model, corpus, **kwargs): # Run standard training model.train(corpus, **kwargs) # Restore original vectors for idx, word in enumerate(model.wv.index_to_key): if word in old_vocab: model.wv.vectors[idx] = old_vectors[model.wv.key_to_index[word]] # Train with custom logic custom_train( original_model, new_corpus, total_examples=len(new_corpus), epochs=5, compute_loss=True ) # Save updated model original_model.save("path/to/updated/model")
Notes:
- This is a bit of a hack, but it works for most cases.
- Be cautious with large models: restoring vectors after each batch adds overhead.
- For better control, you could modify Gensim's internal training loop to only update new word parameters, but that requires diving into the source code.
3. Train a Submodel and Align to Original Space
If your new dataset has enough overlap with the original vocabulary (shared words), you can train a small model on the new data then align its vectors to the original model's space before merging.
How it works:
- Train a small Word2Vec model on your new dataset (include both new words and any original words present in the new data).
- Use Procrustes analysis to align the submodel's vectors to the original model's space using shared words as anchors.
- Extract the aligned new word vectors and add them to the original model.
Example Code Snippet:
import numpy as np from gensim.models import Word2Vec # Load original model original_model = Word2Vec.load("path/to/your/model") original_wv = original_model.wv # Train submodel on new dataset new_corpus = [["newword", "oldword1"], ["oldword2", "newword"]] sub_model = Word2Vec( sentences=new_corpus, vector_size=original_wv.vector_size, window=original_model.window, min_count=1, epochs=10 ) sub_wv = sub_model.wv # Find shared words between models shared_words = [word for word in sub_wv.key_to_index if word in original_wv.key_to_index] if not shared_words: raise ValueError("No shared words to align models!") # Get vectors for shared words original_vecs = original_wv[shared_words] sub_vecs = sub_wv[shared_words] # Align submodel vectors to original space using Procrustes def procrustes_align(source_vecs, target_vecs): # Center both sets of vectors source_centered = source_vecs - np.mean(source_vecs, axis=0) target_centered = target_vecs - np.mean(target_vecs, axis=0) # Compute SVD for alignment U, _, Vt = np.linalg.svd(target_centered.T @ source_centered) rotation = U @ Vt # Align source vectors aligned_vecs = source_centered @ rotation.T + np.mean(target_vecs, axis=0) return aligned_vecs # Align submodel vectors aligned_sub_vecs = procrustes_align(sub_vecs, original_vecs) # Update submodel's shared word vectors to aligned versions for i, word in enumerate(shared_words): sub_wv[word] = aligned_sub_vecs[i] # Extract new words (not in original model) new_words = [word for word in sub_wv.key_to_index if word not in original_wv.key_to_index] new_vecs = sub_wv[new_words] # Add new words to original model original_wv.add_vectors(new_words, new_vecs) # Save updated model original_model.save("path/to/updated/model")
Notes:
- This method ensures new vectors are in the same semantic space as the original model.
- Requires at least a few shared words between the new dataset and original vocabulary to align correctly.
内容的提问来源于stack exchange,提问作者Wargream

