如何在Vespa查询阶段高效将文档关键词纳入排序逻辑?
解决Vespa中关键词反向匹配与多语言分词问题
场景背景
基于带有专家及GPT整理关键词字段的分类文档,匹配长度从单个词到中等段落的查询文本,需优先依据关键词字段返回最适配的分类。现有Schema、排名配置及查询语句如下:
现有Schema配置
document category { field id type int { indexing: summary | attribute } field title type string { indexing: summary | attribute | index index: enable-bm25 } field keywords type array<string> { indexing: summary | index } } field title_embedding type tensor<bfloat16>(x[384]) { indexing: input title | embed bert | attribute | index attribute { distance-metric: angular } } fieldset default { fields: title }
现有排名配置
rank-profile bm25_semantic inherits default { inputs { query(query_embedding) tensor<bfloat16>(x[384]) } first-phase { expression: bm25(title) + matches(keywords) + closeness(field, title_embedding) } }
现有查询语句
SELECT * FROM category WHERE userQuery() OR rank(keywords contains 'X' OR keywords contains 'Y') OR ({targetHits: 100}nearestNeighbor(title_embedding,query_embedding))
当前存在的问题
- 需手动拆分查询文本中的每个token,重复编写
keywords contains 'X'逻辑,扩展性极差 - 多语言(如中文)场景下,基础分词效果无法满足需求
解决方案
一、调整Schema字段配置
1. 优化keywords字段的索引与分词
修改keywords字段配置,添加多语言分词器(以中文jieba为例),同时配置合适的索引类型确保匹配效率:
document category { field id type int { indexing: summary | attribute } field title type string { indexing: summary | attribute | index index: enable-bm25 } field keywords type array<string> { indexing: summary | index index: exact | enable-bm25 # exact确保单个关键词精确匹配,enable-bm25支持分词后的短语匹配 tokenizer: jieba # 中文场景用jieba,英文用default,多语言混合用multilingual分词器 } } field title_embedding type tensor<bfloat16>(x[384]) { indexing: input title | embed bert | attribute | index attribute { distance-metric: angular } } fieldset default { fields: title, keywords # 将keywords加入默认字段集,让userQuery()可同时匹配title和keywords }
二、优化排名配置
强化关键词匹配的权重,确保优先返回关键词命中的文档:
rank-profile bm25_semantic inherits default { inputs { query(query_embedding) tensor<bfloat16>(x[384]) } first-phase { # 提升关键词匹配权重,优先保障关键词命中的文档排名 expression: 2 * matches(keywords) + bm25(title) + closeness(field, title_embedding) } }
三、简化查询语句
无需手动拆分查询token,直接利用Vespa的query算子匹配关键词字段:
SELECT * FROM category WHERE userQuery() OR # 匹配title和keywords(默认字段集包含两者) ({targetHits: 100}nearestNeighbor(title_embedding, query_embedding))
若需明确区分title与keywords的匹配逻辑,可写成:
SELECT * FROM category WHERE userQuery(field=title) OR query(field=keywords) OR ({targetHits: 100}nearestNeighbor(title_embedding, query_embedding))
四、多语言分词适配
- 字段侧:根据文档关键词的语言类型,为
keywords字段配置对应分词器,例如中文用jieba、英文用default、多语言混合用multilingual - 查询侧:若查询语言与字段分词器不一致,可在查询时指定分词器:
SELECT * FROM category WHERE query(field=keywords, tokenizer=jieba)
内容的提问来源于stack exchange,提问作者standalone_2045
相关产品推荐
相关产品推荐

