sklearn的CountVectorizer生成稀疏矩阵后调用get_feature_names()报错如何解决?
问题原因
CountVectorizer的fit_transform()方法返回值为转换后的稀疏词频矩阵,属于scipy的稀疏矩阵类,本身不存储特征名称相关属性,也没有get_feature_names方法- 特征名存储在
CountVectorizer的实例对象中,你代码中错误地对稀疏矩阵调用了该方法,因此触发AttributeError
修正代码
将最后一行的调用对象改为你创建的CountVectorizer实例count_vec即可:
from sklearn.feature_extraction.text import CountVectorizer count_vec = CountVectorizer(ngram_range=(1,2)) bigram_count = count_vec.fit_transform(data["CleanedText"].values) # 适配不同scikit-learn版本 if hasattr(count_vec, 'get_feature_names_out'): # scikit-learn 1.0及以上版本用该方法 feature_names = count_vec.get_feature_names_out() else: # 旧版本用该方法 feature_names = count_vec.get_feature_names() print(feature_names)
扩展使用场景
如果需要把稀疏矩阵转换为带特征列名的pandas DataFrame,可以使用以下代码:
import pandas as pd count_df = pd.DataFrame( bigram_count.toarray(), columns=feature_names )
内容的提问来源于stack exchange,提问作者Yatin Kumar
相关产品推荐
相关产品推荐

