GPU环境下spaCy+Numpy求相似词向量遇类型错误的解决咨询
解决GPU环境下spaCy+NumPy查找相似词向量的TypeError问题
在GPU环境中运行spaCy查找相似词向量的代码时,触发了TypeError,错误提示为:
Implicit conversion to a NumPy array is not allowed. Please use
.get()to construct a NumPy array explicitly
原代码
import numpy as np your_word = "country" ms = nlp.vocab.vectors.most_similar( np.asarray([nlp.vocab.vectors[nlp.vocab.strings[your_word]]]), n=10, ) words = [nlp.vocab.strings[w] for w in ms[0][0]] distances = ms[2] print(words)
报错栈
--------------------------------------------------------------------------- TypeError Traceback (most recent call last) Cell In[139], line 6 1 import numpy as np 3 your_word = "country" 5 ms = nlp.vocab.vectors.most_similar( ----> 6 np.asarray([nlp.vocab.vectors[nlp.vocab.strings[your_word]]]), 7 n=10, 8 ) 10 words = [nlp.vocab.strings[w] for w in ms[0][0]] 11 distances = ms[2] File cupy/_core/core.pyx:1475, in cupy._core.core._ndarray_base.__array__() TypeError: Implicit conversion to a NumPy array is not allowed. Please use `.get()` to construct a NumPy array explicitly.
问题原因
GPU环境下,spaCy的词向量存储在CuPy数组中,直接用np.asarray()尝试将其转为NumPy数组会触发隐式转换限制,CuPy不允许这种隐式操作,必须显式调用*.get()*方法将CuPy数组转为NumPy数组。
修复方案
方案1:显式转为NumPy数组(CPU运算)
import numpy as np your_word = "country" # 显式用.get()将CuPy数组转为NumPy数组 word_vector = nlp.vocab.vectors[nlp.vocab.strings[your_word]].get() ms = nlp.vocab.vectors.most_similar( np.asarray([word_vector]), n=10, ) words = [nlp.vocab.strings[w] for w in ms[0][0]] distances = ms[2] print(words)
方案2:直接使用CuPy数组(GPU加速,推荐)
避免数据回传CPU,保持GPU运算效率:
your_word = "country" word_vector = nlp.vocab.vectors[nlp.vocab.strings[your_word]] # 将向量转为(1, 维度)的形状,满足most_similar的输入要求 ms = nlp.vocab.vectors.most_similar( word_vector.reshape(1, -1), n=10, ) words = [nlp.vocab.strings[w] for w in ms[0][0]] distances = ms[2] print(words)
内容的提问来源于stack exchange,提问作者onder
相关产品推荐
相关产品推荐

