You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Vespa中非英语字段词序保留及排序规则配置问询

Vespa Schema 实现方案

核心Schema定义

针对你的需求,以下是包含藏文威利转写字段的Schema示例,已配置好分词规则和排序优先级:

schema my_document {
    document my_document {
        # 藏文威利转写字段1
        field tibetan_field1 type string {
            indexing: summary | index
            index: enable-bm25
            analyzer: wylie_analyzer
        }
        # 藏文威利转写字段2
        field tibetan_field2 type string {
            indexing: summary | index
            index: enable-bm25
            analyzer: wylie_analyzer
        }
        # 英文主字段示例
        field english_content type string {
            indexing: summary | index
            index: enable-bm25
        }
    }

    # 自定义分词器:保留撇号、加号,关闭词干提取
    analyzer wylie_analyzer {
        tokenizer unicode {
            keep: ['\'', '+']
        }
        stemmer none
        # 若不需要小写转换可移除以下行
        normalizer lowercase
    }

    # 排序优先级配置
    rank-profile default {
        first-phase {
            expression:
                # 优先级1:字段开头精确匹配所有查询token(词序一致)
                if (match(tibetan_field1, "prefix") AND all(tibetan_field1)) then 100.0
                else if (match(tibetan_field2, "prefix") AND all(tibetan_field2)) then 90.0
                # 优先级2:任意位置精确匹配所有查询token(词序一致)
                else if (match(tibetan_field1, "exact") AND all(tibetan_field1)) then 80.0
                else if (match(tibetan_field2, "exact") AND all(tibetan_field2)) then 70.0
                # 优先级3:降级到BM25综合评分
                else bm25(tibetan_field1) + bm25(tibetan_field2) + bm25(english_content);
        }
    }
}

关键配置说明

  • 分词器设置:wylie_analyzer用unicode分词器保留撇号(')和加号(+),stemmer none彻底关闭词干提取,避免藏文转写字符被篡改。
  • 匹配逻辑:
    • match(field, "prefix") 判断查询短语是否匹配字段开头,结合all(field)确保所有查询token都命中且词序一致。
    • match(field, "exact") 判断查询短语是否作为完整序列出现在字段任意位置,同样用all(field)保证全token匹配。
  • 排序权重:通过阶梯式的if-else逻辑给不同匹配场景分配权重,确保开头匹配>任意位置精确匹配>BM25的优先级顺序。

查询效果示例

当查询jam dpal smra ba'i时:

  • 首先返回tibetan_field1开头完全匹配该短语的文档,其次是tibetan_field2开头匹配的文档。
  • 若没有开头匹配的结果,会返回任意位置包含该完整短语的文档。
  • 最后返回BM25评分最高的匹配文档作为兜底。

内容的提问来源于stack exchange,提问作者Roope K

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 00:50:39