如何获取TfidfVectorizer生成特征词的对应词频?
获取Tfidf特征词的词频方法
TfidfVectorizer本身只计算TF-IDF权重,没有直接存储原始词频,但可以通过以下方式获取特征词的词频:
方法一:用CountVectorizer同步统计(推荐)
只要让CountVectorizer和TfidfVectorizer的预处理参数(分词、停用词、大小写转换等)保持一致,两者的特征词列表就完全匹配,直接用CountVectorizer统计词频即可:
from sklearn.feature_extraction.text import TfidfVectorizer, CountVectorizer corpus = [ 'This is the first document.', 'This document is the second document.', 'And this is the third one.', 'Is this the first document?', ] # 初始化两个向量器,参数保持一致 tfidf_vec = TfidfVectorizer() count_vec = CountVectorizer() # 分别拟合语料 tfidf_vec.fit_transform(corpus) count_matrix = count_vec.fit_transform(corpus) # 获取特征词列表 feature_words = tfidf_vec.get_feature_names_out() # 统计每个词的总出现次数 total_counts = count_matrix.sum(axis=0).A1 # 打包成字典,方便查询 word_freq = dict(zip(feature_words, total_counts)) # 打印结果 for word, freq in word_freq.items(): print(f"{word}: {freq}")
运行后会输出每个特征词在所有文档中的总出现次数,比如document会显示5次,完全匹配语料中的实际出现情况。
方法二:通过TfidfVectorizer倒推(不推荐)
TF-IDF的计算公式是tf * idf,其中的tf是经过归一化处理的(默认L2归一化),要还原原始词频需要先取消归一化,再结合idf值计算,步骤繁琐且容易出错,因此更推荐用第一种方法。
内容的提问来源于stack exchange,提问作者james pow
相关产品推荐
相关产品推荐

