使用LatentDirichletAllocation时遇参数错误:__init__不识别n_components
n_components参数错误 首先,先看你遇到的错误栈:
TypeError Traceback (most recent call last)
in ()
23 # tfidf = vectorizer.fit_transform(line)
24 # print(tfidf)
---> 25 lda = LatentDirichletAllocation(n_components = 100)
26 lda.fit(bag_of_words)
27 tf_feature_names = vector.get_feature_names()
TypeError: init() got an unexpected keyword argument 'n_components'
这个错误的核心原因是scikit-learn版本差异导致的参数命名变化,我来给你详细说下怎么解决:
问题根源
在scikit-learn 0.20版本之前,LatentDirichletAllocation类用来定义主题数量的参数是n_topics;而从0.20版本开始,这个参数被重命名为n_components(和其他降维/主题模型的参数命名保持一致)。如果你当前使用的是旧版本的scikit-learn,自然会因为找不到n_components这个参数而报错。
两种解决方案
方案1:适配旧版本,修改参数名
直接把代码里的n_components替换成n_topics即可,修改后的代码如下:
lda = LatentDirichletAllocation(n_topics = 100) lda.fit(bag_of_words)
方案2:升级scikit-learn到新版本
如果你更习惯使用n_components这个参数名,可以通过pip升级你的scikit-learn库:
pip install --upgrade scikit-learn
升级完成后,你的原始代码就能正常运行了。
可选:确认当前scikit-learn版本
你可以先运行下面的代码查看当前版本,确认是不是版本问题:
import sklearn print(sklearn.__version__)
如果版本号小于0.20,那肯定就是这个原因导致的错误啦。
内容的提问来源于stack exchange,提问作者user7120305

