You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于训练好的Gensim Word2Vec模型添加新词并保留旧词向量?

Great question! I’ve run into this exact scenario before—wanting to extend a pre-trained Word2Vec model with new words without messing up the existing embeddings. Here are a few practical approaches that align with your request for a doc2vec-infer-like solution:

1. Infer New Word Vectors Using Pre-Trained Context (Most Similar to Doc2Vec Infer)

This method mimics doc2vec's inference process: we fix the existing word vectors and optimize a new vector for each unseen word until it fits naturally with its context from your new dataset.

How it works:

  • For each new word in your dataset, collect all its context windows (match the window size used to train your original Word2Vec model).
  • Initialize a random vector for the new word.
  • Use the original model's training objective (CBOW or Skip-gram) to tweak this vector:
    • For CBOW: Adjust the new word's vector so that it predicts the average of its context vectors.
    • For Skip-gram: Adjust the new word's vector so that its context words are predicted from it.
  • Stop when the vector stabilizes (or after a fixed number of epochs, like doc2vec's infer_epochs).

Example Code Snippet:

import numpy as np
from gensim.models import Word2Vec

# Load your pre-trained model
original_model = Word2Vec.load("path/to/your/model")
wv = original_model.wv
vector_size = wv.vector_size
window = original_model.window
learning_rate = 0.01  # Adjust based on your needs

def infer_new_word_vector(new_word, context_windows):
    # Initialize random vector
    new_vec = np.random.randn(vector_size) * 0.01
    
    for _ in range(10):  # Mimic infer_epochs
        total_loss = 0.0
        for context in context_windows:
            # Filter out words not in original model (we only trust pre-trained vectors)
            valid_context = [word for word in context if word in wv]
            if not valid_context:
                continue
                
            # CBOW objective: new_vec should predict the average context vector
            target_vec = np.mean(wv[valid_context], axis=0)
            # Calculate error
            error = target_vec - new_vec
            # Update new_vec
            new_vec += learning_rate * error
            total_loss += np.sum(error**2)
        
        if total_loss < 1e-4:  # Stop early if loss is small
            break
    
    return new_vec

# Usage: Collect context windows for your new word
new_word = "newword"
context_windows = [["oldword1", "oldword2"], ["oldword3", "newword", "oldword4"]]  # Example contexts
new_vec = infer_new_word_vector(new_word, context_windows)

# Add the new word to the model's KeyedVectors
wv.add_vectors([new_word], [new_vec])
# Save the updated model
original_model.save("path/to/updated/model")

Notes:

  • This preserves all original embeddings since we never modify them.
  • The quality depends heavily on how many valid context windows you have for each new word (more contexts = better vectors).
  • You can adapt this for Skip-gram by reversing the objective (predict context words from the new word vector).

2. Incremental Training with Frozen Original Embeddings

Gensim's default update=True mode will adjust all vectors (old and new), but we can hack the model to freeze original words while training new ones.

How it works:

  1. Load your pre-trained model and extend its vocabulary with the new dataset.
  2. Freeze the vectors of existing words by marking them as "untrainable" (we'll modify the model's internal weights to prevent updates).
  3. Train only on the new dataset, so only new word vectors are updated.

Example Code Snippet:

from gensim.models import Word2Vec

# Load pre-trained model
original_model = Word2Vec.load("path/to/your/model")
old_vocab = set(original_model.wv.key_to_index.keys())

# Extend vocabulary with new corpus
new_corpus = [["newword", "oldword1"], ["oldword2", "newword"]]  # Your new dataset
original_model.build_vocab(new_corpus, update=True)

# Freeze original word vectors: save old vectors, then restore them after each training batch
old_vectors = original_model.wv.vectors.copy()

# Override the train method to restore old vectors after each update (simplified approach)
def custom_train(model, corpus, **kwargs):
    # Run standard training
    model.train(corpus, **kwargs)
    # Restore original vectors
    for idx, word in enumerate(model.wv.index_to_key):
        if word in old_vocab:
            model.wv.vectors[idx] = old_vectors[model.wv.key_to_index[word]]

# Train with custom logic
custom_train(
    original_model,
    new_corpus,
    total_examples=len(new_corpus),
    epochs=5,
    compute_loss=True
)

# Save updated model
original_model.save("path/to/updated/model")

Notes:

  • This is a bit of a hack, but it works for most cases.
  • Be cautious with large models: restoring vectors after each batch adds overhead.
  • For better control, you could modify Gensim's internal training loop to only update new word parameters, but that requires diving into the source code.

3. Train a Submodel and Align to Original Space

If your new dataset has enough overlap with the original vocabulary (shared words), you can train a small model on the new data then align its vectors to the original model's space before merging.

How it works:

  1. Train a small Word2Vec model on your new dataset (include both new words and any original words present in the new data).
  2. Use Procrustes analysis to align the submodel's vectors to the original model's space using shared words as anchors.
  3. Extract the aligned new word vectors and add them to the original model.

Example Code Snippet:

import numpy as np
from gensim.models import Word2Vec

# Load original model
original_model = Word2Vec.load("path/to/your/model")
original_wv = original_model.wv

# Train submodel on new dataset
new_corpus = [["newword", "oldword1"], ["oldword2", "newword"]]
sub_model = Word2Vec(
    sentences=new_corpus,
    vector_size=original_wv.vector_size,
    window=original_model.window,
    min_count=1,
    epochs=10
)
sub_wv = sub_model.wv

# Find shared words between models
shared_words = [word for word in sub_wv.key_to_index if word in original_wv.key_to_index]
if not shared_words:
    raise ValueError("No shared words to align models!")

# Get vectors for shared words
original_vecs = original_wv[shared_words]
sub_vecs = sub_wv[shared_words]

# Align submodel vectors to original space using Procrustes
def procrustes_align(source_vecs, target_vecs):
    # Center both sets of vectors
    source_centered = source_vecs - np.mean(source_vecs, axis=0)
    target_centered = target_vecs - np.mean(target_vecs, axis=0)
    
    # Compute SVD for alignment
    U, _, Vt = np.linalg.svd(target_centered.T @ source_centered)
    rotation = U @ Vt
    
    # Align source vectors
    aligned_vecs = source_centered @ rotation.T + np.mean(target_vecs, axis=0)
    return aligned_vecs

# Align submodel vectors
aligned_sub_vecs = procrustes_align(sub_vecs, original_vecs)

# Update submodel's shared word vectors to aligned versions
for i, word in enumerate(shared_words):
    sub_wv[word] = aligned_sub_vecs[i]

# Extract new words (not in original model)
new_words = [word for word in sub_wv.key_to_index if word not in original_wv.key_to_index]
new_vecs = sub_wv[new_words]

# Add new words to original model
original_wv.add_vectors(new_words, new_vecs)

# Save updated model
original_model.save("path/to/updated/model")

Notes:

  • This method ensures new vectors are in the same semantic space as the original model.
  • Requires at least a few shared words between the new dataset and original vocabulary to align correctly.

内容的提问来源于stack exchange,提问作者Wargream

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 09:02:48