You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

sklearn的CountVectorizer生成稀疏矩阵后调用get_feature_names()报错如何解决?

问题原因
  • CountVectorizer的fit_transform()方法返回值为转换后的稀疏词频矩阵,属于scipy的稀疏矩阵类,本身不存储特征名称相关属性,也没有get_feature_names方法
  • 特征名存储在CountVectorizer的实例对象中,你代码中错误地对稀疏矩阵调用了该方法,因此触发AttributeError
修正代码

将最后一行的调用对象改为你创建的CountVectorizer实例count_vec即可:

from sklearn.feature_extraction.text import CountVectorizer
count_vec = CountVectorizer(ngram_range=(1,2))
bigram_count = count_vec.fit_transform(data["CleanedText"].values)
# 适配不同scikit-learn版本
if hasattr(count_vec, 'get_feature_names_out'):
    # scikit-learn 1.0及以上版本用该方法
    feature_names = count_vec.get_feature_names_out()
else:
    # 旧版本用该方法
    feature_names = count_vec.get_feature_names()
print(feature_names)
扩展使用场景

如果需要把稀疏矩阵转换为带特征列名的pandas DataFrame,可以使用以下代码:

import pandas as pd
count_df = pd.DataFrame(
    bigram_count.toarray(),
    columns=feature_names
)

内容的提问来源于stack exchange,提问作者Yatin Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 16:54:03