AttributeError问题:CountVectorizer对象无get_feature_names属性
解决AttributeError: 'CountVectorizer' object has no attribute 'get_feature_names'
这个问题是因为scikit-learn版本更新导致的:从1.0版本开始,get_feature_names()方法被弃用,替换为get_feature_names_out()。
解决步骤:
- 先确认你的scikit-learn版本:
import sklearn print(sklearn.__version__) - 如果版本≥1.0,直接把代码中的
model.get_feature_names()替换成model.get_feature_names_out()。 - 要是需要兼容新旧版本,可以用如下兼容写法:
try: w = model.get_feature_names_out() except AttributeError: w = model.get_feature_names()
修改后的完整代码:
from sklearn.feature_extraction.text import CountVectorizer from sklearn.linear_model import LogisticRegression from sklearn.model_selection import train_test_split import pandas as pd c = CountVectorizer(stop_words = 'english') def text_fit(X, y, model, clf_model, coef_show=1): X_c = model.fit_transform(X) print('# features: {}'.format(X_c.shape[1])) X_train, X_test, y_train, y_test = train_test_split(X_c, y, random_state=0) print('# train records: {}'.format(X_train.shape[0])) print('# test records: {}'.format(X_test.shape[0])) clf = clf_model.fit(X_train, y_train) acc = clf.score(X_test, y_test) print ('Model Accuracy: {}'.format(acc)) if coef_show == 1: # 替换为兼容版本的方法 try: w = model.get_feature_names_out() except AttributeError: w = model.get_feature_names() coef = clf.coef_.tolist()[0] coeff_df = pd.DataFrame({'Word' : w, 'Coefficient' : coef}) coeff_df = coeff_df.sort_values(['Coefficient', 'Word'], ascending=[0, 1]) print('') print('-Top 20 positive-') print(coeff_df.head(20).to_string(index=False)) print('') print('-Top 20 negative-') print(coeff_df.tail(20).to_string(index=False)) text_fit(X, y, c, LogisticRegression())
补充说明:
你之前代码正常后来出错,大概率是重建项目时安装的scikit-learn版本比原来的高,新版本移除了旧方法导致报错。用上面的兼容写法可以同时适配新旧版本,避免环境变动引发的问题。
内容的提问来源于stack exchange,提问作者DeValor12
相关产品推荐
相关产品推荐

