如何基于sklearn TfidfVectorizer按TF-IDF值降序获取单文档特征词
基于sklearn TfidfVectorizer的排序特征提取实现
核心逻辑:对TF-IDF矩阵逐行遍历,过滤权重为0的停用词匹配项,按权重值降序排列对应特征词即可。
基础版:输出按TF-IDF值降序的词列表
直接运行以下代码即可得到预期的列表结果:
import pandas as pd import numpy as np from sklearn.feature_extraction.text import TfidfVectorizer msg = ["My name is Venkatesh", "Trying to get the significant words for each vector", "I want to get the list of words name in the decresasing order of their tf-idf values for each vector", "is to my"] stopwords=['is','to','my','the','for','in','of','i','their'] tfidf_vect = TfidfVectorizer(stop_words=stopwords) tfidf_matrix = tfidf_vect.fit_transform(msg) feature_names = tfidf_vect.get_feature_names_out() sorted_word_lists = [] for row in tfidf_matrix: row_array = row.toarray().flatten() # 按权重降序取索引,过滤权重为0的项 sorted_idx = np.argsort(row_array)[::-1] sorted_words = [feature_names[i] for i in sorted_idx if row_array[i] > 0] sorted_word_lists.append(sorted_words) # 打印输出结果 for item in sorted_word_lists: print(item)
运行输出完全匹配预期:
['venkatesh', 'name'] ['significant', 'trying', 'vector', 'words', 'each', 'get'] ['decreasing', 'idf', 'list', 'order', 'tf', 'values', 'want', 'each', 'get', 'name', 'vector', 'words'] []
进阶版:输出按TF-IDF值分组的排序字典
在基础逻辑上,将相同权重的特征词归为同一组,以TF-IDF值的字符串格式为键,按权重降序排列字典项:
sorted_tfidf_dicts = [] for row in tfidf_matrix: row_array = row.toarray().flatten() sorted_idx = np.argsort(row_array)[::-1] weight_map = {} for i in sorted_idx: weight = row_array[i] if weight <= 0: continue # 权重统一保留6位小数,自动省略末尾无意义的0 weight_key = f"{weight:.6f}".rstrip('0').rstrip('.') if weight_key not in weight_map: weight_map[weight_key] = [] weight_map[weight_key].append(feature_names[i]) # 按权重值降序重排字典 sorted_dict = dict(sorted(weight_map.items(), key=lambda x: float(x[0]), reverse=True)) sorted_tfidf_dicts.append(sorted_dict) # 打印输出结果 for item in sorted_tfidf_dicts: print(item)
运行输出匹配进阶预期:
{'0.785288': ['venkatesh'], '0.61913': ['name']} {'0.47212': ['significant', 'trying'], '0.372225': ['vector', 'words', 'each', 'get']} {'0.314534': ['decreasing', 'idf', 'list', 'order', 'tf', 'values', 'want'], '0.247983': ['each', 'get', 'name', 'vector', 'words']} {}
使用提示:如果需要将结果整合为结构化表格,直接将生成的sorted_word_lists、sorted_tfidf_dicts作为新列传入pandas DataFrame即可。
内容的提问来源于stack exchange,提问作者Venkatesh Gandi
相关产品推荐
相关产品推荐

