You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Gensim加载的Google Word2Vec适配自定义词汇并生成嵌入矩阵

How to Populate Your Embedding Matrix with Google Word2Vec for Custom Vocabulary

Great question! Here's a straightforward way to fill your embedding_matrix using the pre-trained Google Word2Vec model, focusing only on your custom vocabulary:

Step-by-Step Implementation

First, ensure you’ve imported numpy (since you’re using np.zeros):

import numpy as np
import gensim

Next, load the model as you already did (note: in newer Gensim versions, load_word2vec_format is part of KeyedVectors, but your existing code works for most versions):

# Load Google's pre-trained Word2Vec model
model = gensim.models.Word2Vec.load_word2vec_format('./model/GoogleNews-vectors-negative300.bin', binary=True)

Now, iterate over your custom vocabulary to fill the embedding matrix. The critical step is checking if each word exists in the pre-trained model’s vocabulary first—this avoids errors for words that aren’t present (like rare terms or domain-specific jargon):

my_vocabulary = ["your", "custom", "words", "here"]  # Replace with your actual vocab list
embedding_matrix = np.zeros((len(my_vocabulary), 300))

for idx, word in enumerate(my_vocabulary):
    # Check if the word exists in the Word2Vec model
    # For Gensim 4.x+ use model.wv.key_to_index; for older versions use model.vocab
    if word in model.wv.key_to_index:
        # Assign the pre-trained vector to the corresponding matrix row
        embedding_matrix[idx] = model.wv[word]
    else:
        # Optional: Handle out-of-vocabulary (OOV) words (pick one option)
        # Option 1: Leave as zero vector (default, since we initialized with zeros)
        pass
        # Option 2: Initialize with random values (often better for model training than zeros)
        # embedding_matrix[idx] = np.random.normal(scale=0.6, size=300)

Key Notes

  • Gensim Version Compatibility: If you’re using an older Gensim version (pre-4.0), swap model.wv.key_to_index with model.vocab in the existence check.
  • OOV Word Handling: Words not found in Google’s model are common. Using random initialization instead of zero vectors typically helps your model learn better representations for these terms during training.
  • Matrix Structure: Each row in embedding_matrix maps directly to the word at the same index in my_vocabulary, so you can easily use this matrix for downstream tasks like neural network embedding layers.

内容的提问来源于stack exchange,提问作者A.Papa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:23:02