You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从Gensim 3迁移至4时出现vocab相关AttributeError问题求助

解决Gensim 3转Gensim 4时的AttributeError问题

问题背景

将Gensim 3版本的Word2Vec代码迁移至Gensim 4版本时,执行X = model[model.wv.vocab]触发AttributeError,原代码用于训练Word2Vec模型,通过PCA降维后绘制词向量散点图,目标是生成词向量的二维散点可视化图。

报错原因

Gensim 4对API做了简化调整:

  • model.wv.vocab不再是可直接用于索引的结构
  • 主模型对象model不再支持通过词汇表对象直接索引词向量
  • 获取词汇表单词列表的方式从list(model.wv.vocab)改为model.wv.index_to_key

修改后的完整代码

from gensim.models import Word2Vec
from sklearn.decomposition import PCA
from matplotlib import pyplot
import numpy as np

# 定义训练数据
sentences = [['this', 'is', 'the', 'first', 'sentence', 'for', 'word2vec'],
            ['this', 'is', 'the', 'second', 'sentence'],
            ['yet', 'another', 'sentence'],
            ['one', 'more', 'sentence'],
            ['and', 'the', 'final', 'sentence']]

# 训练模型
model = Word2Vec(sentences, min_count=1)

# 拟合2D PCA模型到词向量
X = model.wv.vectors  # 直接获取所有词向量矩阵
pca = PCA(n_components=2)
result = pca.fit_transform(X)

# 绘制散点图
pyplot.scatter(result[:, 0], result[:, 1])
words = model.wv.index_to_key  # 获取词汇表单词列表
for i, word in enumerate(words):
    pyplot.annotate(word, xy=(result[i, 0], result[i, 1]))
pyplot.show()

关键修改说明

  1. 替换X = model[model.wv.vocab]为X = model.wv.vectors:model.wv.vectors直接返回所有词向量组成的numpy矩阵,无需通过词汇表索引
  2. 替换words = list(model.wv.vocab)为words = model.wv.index_to_key:index_to_key是Gensim 4中获取词汇表单词列表的标准方式,顺序与vectors矩阵的行一一对应

内容的提问来源于stack exchange,提问作者Manwest

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 01:05:26