You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改searchSimilarDocumentsByPhrases函数实现结果按计数降序?

问题:修改Python函数实现短语出现次数降序排序

我编写了如下Python函数searchSimilarDocumentsByPhrases,用于通过短语搜索相似文档并统计短语在各文档中的出现次数:

def searchSimilarDocumentsByPhrases(corpus, Ids, contractIds,count,phrases=None):
  tfidf = TfidfVectorizer(vocabulary = phrases, ngram_range=(1, 6))
  tfs = tfidf.fit_transform(corpus)
  feature_names = tfidf.get_feature_names_out()
  rows, cols = tfs.nonzero()

  phrase_counts = defaultdict(list)
  for row, col in zip(rows, cols):
    phraseCount = corpus[row].count(feature_names[col])
    phrase_counts[feature_names[col]].append({Ids[row]:{contractIds[row]: phraseCount}})

  counter=count
  phraselist=[]
  for phrase, counts in phrase_counts.items():
    counts.sort(key=lambda x: list(x.items()), reverse=True) 
    counts=counts[:counter]
    phraselist.append(phrase)
    phraselist.append(counts)
  return phraselist

当传入如下输入参数时:

{ "phrases":
 [
  "test and evaluation"
 ],
 "count":"16"
}

当前输出结果未按短语出现次数降序排列,我期望结果按每个条目内的计数值从高到低排序,预期输出示例如下:

[
    "test and evaluation",
    [
        {
            "1080": {
                "LMLB_C-59": 21
            }
        },
        ...
    ]
]

请告知如何修改函数实现该排序需求?


修改方案

问题出在排序的key参数上,当前代码是按文档ID排序,而非短语出现次数。需要调整排序逻辑,提取每个条目里的计数值作为排序依据,同时修正count参数的类型问题。

修改后的完整函数如下:

from collections import defaultdict
from sklearn.feature_extraction.text import TfidfVectorizer

def searchSimilarDocumentsByPhrases(corpus, Ids, contractIds, count, phrases=None):
    tfidf = TfidfVectorizer(vocabulary=phrases, ngram_range=(1, 6))
    tfs = tfidf.fit_transform(corpus)
    feature_names = tfidf.get_feature_names_out()
    rows, cols = tfs.nonzero()

    phrase_counts = defaultdict(list)
    for row, col in zip(rows, cols):
        phraseCount = corpus[row].count(feature_names[col])
        phrase_counts[feature_names[col]].append({Ids[row]: {contractIds[row]: phraseCount}})

    counter = int(count)  # 将字符串类型的count转为整数
    phraselist = []
    for phrase, counts in phrase_counts.items():
        # 提取每个条目内的计数值作为排序依据
        counts.sort(key=lambda x: list(list(x.values())[0].values())[0], reverse=True)
        counts = counts[:counter]
        phraselist.append(phrase)
        phraselist.append(counts)
    return phraselist

关键修改说明

  1. 排序逻辑修正:将原排序代码counts.sort(key=lambda x: list(x.items()), reverse=True)替换为counts.sort(key=lambda x: list(list(x.values())[0].values())[0], reverse=True)。
    • 每个条目是{"文档ID": {"合同ID": 出现次数}}的嵌套字典,通过list(x.values())[0]拿到内层合同字典,再通过list(...values())[0]提取出次数值,以此作为降序排序的依据。
  2. 参数类型转换:输入的count是字符串(如"16"),转为整数int(count)后才能正常执行切片操作。

内容的提问来源于stack exchange,提问作者amnpawar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 03:32:56