使用CountVectorizer触发AttributeError:get_feature_names_out未找到
问题:情感分析代码调用get_feature_names_out触发AttributeError
参考YouTube公开代码自定义情感分析时,调用get_feature_names_out()触发报错,尝试get_features_name也无效。
代码片段
import re import time from tqdm import tqdm import numpy as np import pandas as pd import matplotlib.pyplot as plt import plotly.graph_objects as go from plotly.subplots import make_subplots import nltk from nltk.corpus import stopwords as nltk_stopwords from pymystem3 import Mystem from wordcloud import WordCloud from sklearn.feature_extraction.text import CountVectorizer from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.model_selection import cross_validate from sklearn.model_selection import train_test_split from sklearn.linear_model import LogisticRegression from sklearn.metrics import * app_counter = CountVectorizer(ngram_range=(1, 1)) app_count = app_counter.fit_transform(app_reviews_df['text_clear']) app_count.toarray().sum(axis = 0) app_count.get_feature_names_out().shape
报错信息
--------------------------------------------------------------------------- AttributeError Traceback (most recent call last) <ipython-input-67-5e08d135d61f> in <module> ----> 1 app_count.get_feature_names_out().shape /usr/local/lib/python3.8/dist-packages/scipy/sparse/base.py in __getattr__(self, attr) 685 return self.getnnz() 686 else: --> 687 raise AttributeError(attr + " not found") 688 689 def transpose(self, axes=None, copy=False): AttributeError: get_feature_names_out 未找到
解决方案
- 核心问题:
get_feature_names_out()是CountVectorizer实例(代码中的app_counter)的方法,不是fit_transform()返回的稀疏矩阵(app_count,属于scipy.sparse类型)的方法。你错误地在稀疏矩阵对象上调用了该方法,导致报错。 - 正确调用方式:
- 若使用scikit-learn 1.0及以上版本:调用
app_counter.get_feature_names_out() - 若使用scikit-learn 1.0以下版本:调用
app_counter.get_feature_names()(注意是复数names)
- 若使用scikit-learn 1.0及以上版本:调用
- 修正后的代码片段:
app_counter = CountVectorizer(ngram_range=(1, 1)) app_count = app_counter.fit_transform(app_reviews_df['text_clear']) # 统计词频总和 word_counts = app_count.toarray().sum(axis=0) # 获取特征名称并查看形状 feature_names = app_counter.get_feature_names_out() print(feature_names.shape)
- 补充说明:
fit_transform()返回的稀疏矩阵仅存储文本的词频稀疏表示,特征名称的管理由CountVectorizer实例负责,因此必须从app_counter对象上获取特征名称。
内容的提问来源于stack exchange,提问作者HOTPEPE
相关产品推荐
相关产品推荐

