You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取TFIDF转换器第1001列的对应词汇?

问题解决:获取TFIDF对应列的正确词汇

你用错对象了!TfidfTransformer本身不存储文本特征的词汇表,它只是对CountVectorizer输出的词频矩阵做TF-IDF转换,所以它的get_feature_names_out()返回的是默认占位符(比如x1000这种),不是实际的单词。

正确的做法是从CountVectorizer对象里获取特征名——词汇表是在CountVectorizer拟合数据时生成并保存的,而且TFIDF矩阵的列顺序和CountVectorizer的词频矩阵列顺序完全一致。

修改后的代码:

count_vectorizer = CountVectorizer()
bag_of_words = count_vectorizer.fit_transform(df)

TFIDF_transformer = TfidfTransformer(norm = 'l2')
TFIDF_representation = TFIDF_transformer.fit_transform(bag_of_words)

# 从count_vectorizer提取特征名,索引1000对应第1001列(Python为0起始索引)
count_vectorizer.get_feature_names_out()[1000]

内容的提问来源于stack exchange,提问作者LY T

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 08:54:57