使用Gensim库的Word2Vec无法运行,Scipy导入错误求助
问题
尝试使用Gensim库的Word2Vec模型对数据集进行向量化处理时,运行代码遭遇Scipy抛出的导入错误,代码如下:
from gensim.models import Word2Vec from nltk.tokenize import word_tokenize import nltk nltk.download('punkt') # Sample sentences sentences = [ "This is a sample sentence.", "Word embeddings are cool.", "I love natural language processing." ] # Tokenize the sentences tokenized_sentences = [word_tokenize(sentence.lower()) for sentence in sentences] # Train the Word2Vec model model = Word2Vec(sentences=tokenized_sentences, vector_size=100, window=5, min_count=1, workers=4) # Get the vector representation of a word word_vector = model.wv['sample'] print("Vector representation of 'sample':", word_vector) # Get the most similar words to a given word similar_words = model.wv.most_similar('sample') print("Words most similar to 'sample':", similar_words)
报错信息:
ImportError: cannot import name 'triu' from 'scipy.linalg.special_matrices' (C:\Users\onkar\anaconda3\lib\site-packages\scipy\linalg\special_matrices.py)
已尝试Scipy 1.10、1.12版本及最新版本,问题仍未解决。
解决方法
方法1:安装兼容的Gensim版本
该错误核心是Gensim与Scipy版本不匹配,triu在Scipy的模块位置变更后,部分Gensim版本未同步适配。安装明确兼容的Gensim版本即可解决:
pip install gensim==4.3.2
方法2:手动修改Gensim的导入代码(临时修复)
若不想更换版本,可直接修改Gensim依赖的导入语句:
- 找到本地Python环境中Gensim的安装目录(比如Anaconda环境下的
C:\Users\onkar\anaconda3\lib\site-packages\gensim) - 打开
matutils.py文件,将其中的from scipy.linalg.special_matrices import triu替换为from scipy.linalg import triu - 保存文件后重新运行代码
方法3:替换为其他词向量工具
如果上述方法均无效,可改用以下替代方案:
- spaCy预训练模型:
import spacy # 先安装模型:python -m spacy download en_core_web_sm nlp = spacy.load("en_core_web_sm") doc = nlp("sample") print("Vector representation of 'sample':", doc.vector)
- Transformers预训练模型:
from transformers import BertTokenizer, BertModel import torch # 先安装依赖:pip install transformers torch tokenizer = BertTokenizer.from_pretrained('bert-base-uncased') model = BertModel.from_pretrained('bert-base-uncased') inputs = tokenizer("sample", return_tensors="pt") outputs = model(**inputs) # 取token向量的均值作为词向量 word_vector = outputs.last_hidden_state.mean(dim=1).detach().numpy() print("Vector representation of 'sample':", word_vector)
内容的提问来源于stack exchange,提问作者Onkar Mehra
相关产品推荐
相关产品推荐

