You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于字典列表的文本-分类最佳匹配最优方案咨询

基于字典列表的文本-分类最佳匹配方案

问题说明

需要根据输入文本,从给定的字典列表中匹配最相关的分类,而非严格的完全文本相等。现有代码仅判断文本完全匹配,无法处理类似"computer update"、"update operating system"这类部分匹配的场景。

示例数据

data=[{
"categories":"application update",
"options":["application","update computer application","update computer application","update application computer"],    
},
{
"categories":"computer information",
"options":["computer information","computer information","computer properties"],    
},
{
"categories":"operating system",
"options":["operating system","update computer operating system","computer operating system"],    
},
{
"categories":"adobe software",
"options":["adobe software issue","application adobe issue","fix my adobe application"],    
},
]

目标文本示例

  • text="computer update"
  • text="update operating system"(应匹配operating system分类)
  • text="fix adobe software"(应匹配adobe software分类)

现有无效代码

for i in data:
  for j in i['options']:
    if text== j:
      print(i['categories'])

解决方案:关键词匹配得分法

通过计算输入文本与分类选项的单词匹配度,得分最高的分类即为最佳匹配。核心思路是:将文本拆分为单词集合,统计交集数量作为匹配得分,最终取最高分对应的分类。

基础实现代码

def get_best_category(text, data):
    max_score = -1
    best_category = None
    # 统一转小写并拆分单词,避免大小写干扰
    text_words = set(text.lower().split())
    
    for item in data:
        current_score = 0
        # 遍历当前分类的所有选项,累加匹配得分
        for option in item['options']:
            option_words = set(option.lower().split())
            # 计算两个单词集合的交集大小,作为单条选项的匹配度
            current_score += len(text_words & option_words)
        
        # 更新最高得分与对应分类
        if current_score > max_score:
            max_score = current_score
            best_category = item['categories']
    
    return best_category

# 测试示例
print(get_best_category("computer update", data))          # 输出: application update
print(get_best_category("update operating system", data))  # 输出: operating system
print(get_best_category("fix adobe software", data))       # 输出: adobe software

优化版(去重选项提升效率)

原数据中部分选项存在重复,可先去重减少冗余计算:

def get_best_category(text, data):
    max_score = -1
    best_category = None
    text_words = set(text.lower().split())
    
    for item in data:
        current_score = 0
        # 对选项去重,降低计算量
        unique_options = list(set(item['options']))
        for option in unique_options:
            option_words = set(option.lower().split())
            current_score += len(text_words & option_words)
        
        if current_score > max_score:
            max_score = current_score
            best_category = item['categories']
    
    return best_category

内容的提问来源于stack exchange,提问作者new_dev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 02:17:35