You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LDA主题建模报错:'CountVectorizer'对象无'get_feature_names'属性

解决pyLDAvis调用时CountVectorizer无get_feature_names属性的问题

问题根源

scikit-learn在1.0版本后移除了CountVectorizer的get_feature_names()方法,替换为get_feature_names_out()。如果你的sklearn版本是1.0+,而pyLDAvis版本未适配该变更,就会触发这个报错。

解决方案

1. 手动传入特征名(推荐)

直接调用get_feature_names_out()获取特征名,手动传给pyLDAvis.sklearn.prepare()的feature_names参数,绕过内部旧方法调用:

# 假设你的变量定义如下
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.decomposition import LatentDirichletAllocation
import pyLDAvis.sklearn

vectorizer = CountVectorizer(...)  # 你的CountVectorizer实例
lda = LatentDirichletAllocation(...)  # 训练好的LDA模型
doc_term_matrix = vectorizer.fit_transform(corpus)  # 词频矩阵

# 手动提取特征名并传入
feature_names = vectorizer.get_feature_names_out()
vis = pyLDAvis.sklearn.prepare(lda, doc_term_matrix, vectorizer, feature_names=feature_names)

2. 更新pyLDAvis到最新版

部分新版本的pyLDAvis已适配sklearn的API变更,执行以下命令更新后再尝试原代码:

pip install --upgrade pyLDAvis

3. 降级scikit-learn到兼容版本

若不想修改代码,可以降级sklearn到1.0以下版本(比如0.24.2):

pip install scikit-learn==0.24.2

注意:此方法不推荐长期使用,旧版本可能存在安全或功能缺陷。

调试验证步骤

  • 先确认sklearn版本:
    import sklearn
    print(sklearn.__version__)
    
  • 确认vectorizer确实是CountVectorizer实例:
    print(type(vectorizer))  # 应输出 <class 'sklearn.feature_extraction.text.CountVectorizer'>
    

内容的提问来源于stack exchange,提问作者ASH

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 19:32:51