You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Gensim库的Word2Vec无法运行,Scipy导入错误求助

问题

尝试使用Gensim库的Word2Vec模型对数据集进行向量化处理时,运行代码遭遇Scipy抛出的导入错误,代码如下:

from gensim.models import Word2Vec
from nltk.tokenize import word_tokenize
import nltk
nltk.download('punkt')

# Sample sentences
sentences = [
    "This is a sample sentence.",
    "Word embeddings are cool.",
    "I love natural language processing."
]

# Tokenize the sentences
tokenized_sentences = [word_tokenize(sentence.lower()) for sentence in sentences]

# Train the Word2Vec model
model = Word2Vec(sentences=tokenized_sentences, vector_size=100, window=5, min_count=1, workers=4)

# Get the vector representation of a word
word_vector = model.wv['sample']
print("Vector representation of 'sample':", word_vector)

# Get the most similar words to a given word
similar_words = model.wv.most_similar('sample')
print("Words most similar to 'sample':", similar_words)

报错信息:

ImportError: cannot import name 'triu' from 'scipy.linalg.special_matrices' (C:\Users\onkar\anaconda3\lib\site-packages\scipy\linalg\special_matrices.py)

已尝试Scipy 1.10、1.12版本及最新版本,问题仍未解决。

解决方法

方法1:安装兼容的Gensim版本

该错误核心是Gensim与Scipy版本不匹配,triu在Scipy的模块位置变更后,部分Gensim版本未同步适配。安装明确兼容的Gensim版本即可解决:

pip install gensim==4.3.2

方法2:手动修改Gensim的导入代码(临时修复)

若不想更换版本,可直接修改Gensim依赖的导入语句:

  1. 找到本地Python环境中Gensim的安装目录(比如Anaconda环境下的C:\Users\onkar\anaconda3\lib\site-packages\gensim)
  2. 打开matutils.py文件,将其中的from scipy.linalg.special_matrices import triu替换为from scipy.linalg import triu
  3. 保存文件后重新运行代码

方法3:替换为其他词向量工具

如果上述方法均无效,可改用以下替代方案:

  • spaCy预训练模型:
import spacy
# 先安装模型:python -m spacy download en_core_web_sm
nlp = spacy.load("en_core_web_sm")
doc = nlp("sample")
print("Vector representation of 'sample':", doc.vector)
  • Transformers预训练模型:
from transformers import BertTokenizer, BertModel
import torch

# 先安装依赖:pip install transformers torch
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
model = BertModel.from_pretrained('bert-base-uncased')

inputs = tokenizer("sample", return_tensors="pt")
outputs = model(**inputs)
# 取token向量的均值作为词向量
word_vector = outputs.last_hidden_state.mean(dim=1).detach().numpy()
print("Vector representation of 'sample':", word_vector)

内容的提问来源于stack exchange,提问作者Onkar Mehra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 00:12:47