You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Word2Vec嵌入矩阵构建报错:KeyError 'https' 问题求助

解决Word2Vec嵌入矩阵构建中的KeyError问题

报错原因是你的分词器(tokenizer)中存在https这类不在预训练Word2Vec模型词汇表里的词,直接通过model_1.wv[word]访问会触发KeyError。以下是几种可行的解决方法:

方法1:使用get方法避免报错

直接用model_1.wv.get(word)替代索引访问,不存在的词会返回None,后续逻辑会跳过这些词,保持嵌入矩阵对应位置为初始的0向量。修改后的代码:

vocab_size = len(tokenizer.word_index)+1
embedding_matrix = np.zeros((vocab_size, embedding_vector_size))
# +1 is done because i starts from 1 instead of 0, and goes till len(vocab)
for word, i in tokenizer.word_index.items():
    embedding_vector = model_1.wv.get(word)  # 用get方法避免KeyError
    if embedding_vector is not None:
        embedding_matrix[i] = embedding_vector

方法2:提前过滤无效词汇

先判断词汇是否存在于Word2Vec模型的词汇表中,仅处理有效词汇:

vocab_size = len(tokenizer.word_index)+1
embedding_matrix = np.zeros((vocab_size, embedding_vector_size))

for word, i in tokenizer.word_index.items():
    # 先验证词汇是否在Word2Vec模型的词汇集合内
    if word in model_1.wv.key_to_index:
        embedding_vector = model_1.wv[word]
        embedding_matrix[i] = embedding_vector

方法3:为未登录词生成替代向量

如果希望未登录词(OOV)也能获得有意义的向量,可以用所有Word2Vec向量的均值作为默认值:

# 计算Word2Vec所有向量的均值,作为未登录词的默认向量
oov_vector = np.mean(model_1.wv.vectors, axis=0)

vocab_size = len(tokenizer.word_index)+1
embedding_matrix = np.zeros((vocab_size, embedding_vector_size))

for word, i in tokenizer.word_index.items():
    try:
        embedding_vector = model_1.wv[word]
    except KeyError:
        embedding_vector = oov_vector  # 用均值向量替代未登录词
    embedding_matrix[i] = embedding_vector

各方法适用场景

  • 方法1:实现最简单,适合对未登录词处理要求较低的场景,直接保留0向量。
  • 方法2:严格控制词汇表,避免无效词汇干扰,适合需要精准词汇映射的场景。
  • 方法3:给未登录词分配语义均值,比0向量更能提供有效信息,适合希望模型对未登录词有一定泛化能力的场景。

内容的提问来源于stack exchange,提问作者uhhh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 20:10:50