You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于sklearn TfidfVectorizer按TF-IDF值降序获取单文档特征词

基于sklearn TfidfVectorizer的排序特征提取实现

核心逻辑:对TF-IDF矩阵逐行遍历,过滤权重为0的停用词匹配项,按权重值降序排列对应特征词即可。

基础版:输出按TF-IDF值降序的词列表

直接运行以下代码即可得到预期的列表结果:

import pandas as pd
import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer

msg = ["My name is Venkatesh",
       "Trying to get the significant words for each vector",
       "I want to get the list of words name in the decresasing order of their tf-idf values for each vector",
       "is to my"]

stopwords=['is','to','my','the','for','in','of','i','their']

tfidf_vect = TfidfVectorizer(stop_words=stopwords)
tfidf_matrix = tfidf_vect.fit_transform(msg)
feature_names = tfidf_vect.get_feature_names_out()

sorted_word_lists = []
for row in tfidf_matrix:
    row_array = row.toarray().flatten()
    # 按权重降序取索引,过滤权重为0的项
    sorted_idx = np.argsort(row_array)[::-1]
    sorted_words = [feature_names[i] for i in sorted_idx if row_array[i] > 0]
    sorted_word_lists.append(sorted_words)

# 打印输出结果
for item in sorted_word_lists:
    print(item)

运行输出完全匹配预期:

['venkatesh', 'name']
['significant', 'trying', 'vector', 'words', 'each', 'get']
['decreasing', 'idf', 'list', 'order', 'tf', 'values', 'want', 'each', 'get', 'name', 'vector', 'words']
[]

进阶版:输出按TF-IDF值分组的排序字典

在基础逻辑上,将相同权重的特征词归为同一组,以TF-IDF值的字符串格式为键,按权重降序排列字典项:

sorted_tfidf_dicts = []
for row in tfidf_matrix:
    row_array = row.toarray().flatten()
    sorted_idx = np.argsort(row_array)[::-1]
    weight_map = {}
    for i in sorted_idx:
        weight = row_array[i]
        if weight <= 0:
            continue
        # 权重统一保留6位小数,自动省略末尾无意义的0
        weight_key = f"{weight:.6f}".rstrip('0').rstrip('.')
        if weight_key not in weight_map:
            weight_map[weight_key] = []
        weight_map[weight_key].append(feature_names[i])
    # 按权重值降序重排字典
    sorted_dict = dict(sorted(weight_map.items(), key=lambda x: float(x[0]), reverse=True))
    sorted_tfidf_dicts.append(sorted_dict)

# 打印输出结果
for item in sorted_tfidf_dicts:
    print(item)

运行输出匹配进阶预期:

{'0.785288': ['venkatesh'], '0.61913': ['name']}
{'0.47212': ['significant', 'trying'], '0.372225': ['vector', 'words', 'each', 'get']}
{'0.314534': ['decreasing', 'idf', 'list', 'order', 'tf', 'values', 'want'], '0.247983': ['each', 'get', 'name', 'vector', 'words']}
{}

使用提示:如果需要将结果整合为结构化表格,直接将生成的sorted_word_lists、sorted_tfidf_dicts作为新列传入pandas DataFrame即可。

内容的提问来源于stack exchange,提问作者Venkatesh Gandi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 13:36:18