You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取TfidfVectorizer生成特征词的对应词频?

获取Tfidf特征词的词频方法

TfidfVectorizer本身只计算TF-IDF权重,没有直接存储原始词频,但可以通过以下方式获取特征词的词频:

方法一:用CountVectorizer同步统计(推荐)

只要让CountVectorizer和TfidfVectorizer的预处理参数(分词、停用词、大小写转换等)保持一致,两者的特征词列表就完全匹配,直接用CountVectorizer统计词频即可:

from sklearn.feature_extraction.text import TfidfVectorizer, CountVectorizer

corpus = [
    'This is the first document.',
    'This document is the second document.',
    'And this is the third one.',
    'Is this the first document?',
]

# 初始化两个向量器,参数保持一致
tfidf_vec = TfidfVectorizer()
count_vec = CountVectorizer()

# 分别拟合语料
tfidf_vec.fit_transform(corpus)
count_matrix = count_vec.fit_transform(corpus)

# 获取特征词列表
feature_words = tfidf_vec.get_feature_names_out()

# 统计每个词的总出现次数
total_counts = count_matrix.sum(axis=0).A1

# 打包成字典,方便查询
word_freq = dict(zip(feature_words, total_counts))

# 打印结果
for word, freq in word_freq.items():
    print(f"{word}: {freq}")

运行后会输出每个特征词在所有文档中的总出现次数,比如document会显示5次,完全匹配语料中的实际出现情况。

方法二:通过TfidfVectorizer倒推(不推荐)

TF-IDF的计算公式是tf * idf,其中的tf是经过归一化处理的(默认L2归一化),要还原原始词频需要先取消归一化,再结合idf值计算,步骤繁琐且容易出错,因此更推荐用第一种方法。

内容的提问来源于stack exchange,提问作者james pow

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 03:35:24