如何修改searchSimilarDocumentsByPhrases函数实现结果按计数降序?
问题:修改Python函数实现短语出现次数降序排序
我编写了如下Python函数searchSimilarDocumentsByPhrases,用于通过短语搜索相似文档并统计短语在各文档中的出现次数:
def searchSimilarDocumentsByPhrases(corpus, Ids, contractIds,count,phrases=None): tfidf = TfidfVectorizer(vocabulary = phrases, ngram_range=(1, 6)) tfs = tfidf.fit_transform(corpus) feature_names = tfidf.get_feature_names_out() rows, cols = tfs.nonzero() phrase_counts = defaultdict(list) for row, col in zip(rows, cols): phraseCount = corpus[row].count(feature_names[col]) phrase_counts[feature_names[col]].append({Ids[row]:{contractIds[row]: phraseCount}}) counter=count phraselist=[] for phrase, counts in phrase_counts.items(): counts.sort(key=lambda x: list(x.items()), reverse=True) counts=counts[:counter] phraselist.append(phrase) phraselist.append(counts) return phraselist
当传入如下输入参数时:
{ "phrases": [ "test and evaluation" ], "count":"16" }
当前输出结果未按短语出现次数降序排列,我期望结果按每个条目内的计数值从高到低排序,预期输出示例如下:
[ "test and evaluation", [ { "1080": { "LMLB_C-59": 21 } }, ... ] ]
请告知如何修改函数实现该排序需求?
修改方案
问题出在排序的key参数上,当前代码是按文档ID排序,而非短语出现次数。需要调整排序逻辑,提取每个条目里的计数值作为排序依据,同时修正count参数的类型问题。
修改后的完整函数如下:
from collections import defaultdict from sklearn.feature_extraction.text import TfidfVectorizer def searchSimilarDocumentsByPhrases(corpus, Ids, contractIds, count, phrases=None): tfidf = TfidfVectorizer(vocabulary=phrases, ngram_range=(1, 6)) tfs = tfidf.fit_transform(corpus) feature_names = tfidf.get_feature_names_out() rows, cols = tfs.nonzero() phrase_counts = defaultdict(list) for row, col in zip(rows, cols): phraseCount = corpus[row].count(feature_names[col]) phrase_counts[feature_names[col]].append({Ids[row]: {contractIds[row]: phraseCount}}) counter = int(count) # 将字符串类型的count转为整数 phraselist = [] for phrase, counts in phrase_counts.items(): # 提取每个条目内的计数值作为排序依据 counts.sort(key=lambda x: list(list(x.values())[0].values())[0], reverse=True) counts = counts[:counter] phraselist.append(phrase) phraselist.append(counts) return phraselist
关键修改说明
- 排序逻辑修正:将原排序代码
counts.sort(key=lambda x: list(x.items()), reverse=True)替换为counts.sort(key=lambda x: list(list(x.values())[0].values())[0], reverse=True)。- 每个条目是
{"文档ID": {"合同ID": 出现次数}}的嵌套字典,通过list(x.values())[0]拿到内层合同字典,再通过list(...values())[0]提取出次数值,以此作为降序排序的依据。
- 每个条目是
- 参数类型转换:输入的
count是字符串(如"16"),转为整数int(count)后才能正常执行切片操作。
内容的提问来源于stack exchange,提问作者amnpawar
相关产品推荐
相关产品推荐

