You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Vespa查询阶段高效将文档关键词纳入排序逻辑?

解决Vespa中关键词反向匹配与多语言分词问题

场景背景

基于带有专家及GPT整理关键词字段的分类文档,匹配长度从单个词到中等段落的查询文本,需优先依据关键词字段返回最适配的分类。现有Schema、排名配置及查询语句如下:

现有Schema配置

document category {
        field id type int {
            indexing: summary | attribute
        }

        field title type string {
            indexing: summary | attribute | index
            index: enable-bm25
        }

        field keywords type array<string> {
            indexing: summary | index
        }
}

field title_embedding type tensor<bfloat16>(x[384]) {
        indexing: input title | embed bert | attribute | index
        attribute {
            distance-metric: angular
        }
    }

fieldset default {
        fields: title
    }

现有排名配置

rank-profile bm25_semantic inherits default {
        inputs {
            query(query_embedding) tensor<bfloat16>(x[384])
        }
        first-phase {
            expression: bm25(title) + matches(keywords) + closeness(field, title_embedding)
        }
    }

现有查询语句

SELECT * FROM category WHERE userQuery() OR rank(keywords contains 'X' OR keywords contains 'Y') OR ({targetHits: 100}nearestNeighbor(title_embedding,query_embedding))

当前存在的问题

  • 需手动拆分查询文本中的每个token,重复编写keywords contains 'X'逻辑,扩展性极差
  • 多语言(如中文)场景下,基础分词效果无法满足需求

解决方案

一、调整Schema字段配置

1. 优化keywords字段的索引与分词

修改keywords字段配置,添加多语言分词器(以中文jieba为例),同时配置合适的索引类型确保匹配效率:

document category {
    field id type int {
        indexing: summary | attribute
    }

    field title type string {
        indexing: summary | attribute | index
        index: enable-bm25
    }

    field keywords type array<string> {
        indexing: summary | index
        index: exact | enable-bm25  # exact确保单个关键词精确匹配,enable-bm25支持分词后的短语匹配
        tokenizer: jieba  # 中文场景用jieba,英文用default,多语言混合用multilingual分词器
    }
}

field title_embedding type tensor<bfloat16>(x[384]) {
    indexing: input title | embed bert | attribute | index
    attribute {
        distance-metric: angular
    }
}

fieldset default {
    fields: title, keywords  # 将keywords加入默认字段集,让userQuery()可同时匹配title和keywords
}

二、优化排名配置

强化关键词匹配的权重,确保优先返回关键词命中的文档:

rank-profile bm25_semantic inherits default {
    inputs {
        query(query_embedding) tensor<bfloat16>(x[384])
    }
    first-phase {
        # 提升关键词匹配权重,优先保障关键词命中的文档排名
        expression: 2 * matches(keywords) + bm25(title) + closeness(field, title_embedding)
    }
}

三、简化查询语句

无需手动拆分查询token,直接利用Vespa的query算子匹配关键词字段:

SELECT * FROM category WHERE 
  userQuery() OR  # 匹配title和keywords(默认字段集包含两者)
  ({targetHits: 100}nearestNeighbor(title_embedding, query_embedding))

若需明确区分title与keywords的匹配逻辑,可写成:

SELECT * FROM category WHERE 
  userQuery(field=title) OR 
  query(field=keywords) OR 
  ({targetHits: 100}nearestNeighbor(title_embedding, query_embedding))

四、多语言分词适配

  • 字段侧:根据文档关键词的语言类型,为keywords字段配置对应分词器,例如中文用jieba、英文用default、多语言混合用multilingual
  • 查询侧:若查询语言与字段分词器不一致,可在查询时指定分词器:
    SELECT * FROM category WHERE query(field=keywords, tokenizer=jieba)
    

内容的提问来源于stack exchange,提问作者standalone_2045

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 06:37:52