You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Milvus标量字段前缀搜索失效问题排查求助

Milvus 2.4.8标量字段前缀搜索失效问题

环境

Ubuntu 22.04 单机版 Milvus 2.4.8

问题现象

标量字段channels的前缀搜索仅对第一个token生效,第二个token的前缀搜索无法命中结果,但中缀搜索可以正常工作,怀疑字段未被正确分词。

集合Schema

{
  "fields": [
    {
      "name": "id",
      "description": "Id field",
      "data_type": "DataType.VarChar",
      "is_primary_key": true,
      "max_length": this.NORMALISED_GUID_LENGTH
    },
    {
      "name": "vector",
      "description": "Vector field",
      "data_type": "DataType.FloatVector",
      "dim": this.embeddingModel.dims
    },
    {
      "name": "tag",
      "description": "The partition tag",
      "data_type": "DataType.VarChar",
      "max_length": this.NORMALISED_TAG_LENGTH
    },
    {
      "name": "channels",
      "description": "The channels that may access this entry",
      "data_type": "DataType.VarChar",
      "max_length": this.MAX_CHANNELS_LENGTH
    },
    {
      "name": "payload",
      "description": "The payload meta data",
      "data_type": "DataType.JSON"
    }
  ],
  "partition_key_field": "tag",
  "index_params": [
    {
      "field_name": "vector",
      "index_type": "DISKANN",
      "metric_type": "IP"
    },
    {
      "field_name": "tag",
      "index_name": "tag_index",
      "index_type": "INVERTED"
    },
    {
      "field_name": "channels",
      "index_name": "channels_index",
      "index_type": "INVERTED"
    }
  ]
}

字段示例值

channels字段的示例内容:

"48d302b1963841c39790fecf56b91ddc c8a80b710e455460ae8b2399f5adfef5 b9e8870f9f994bc6a172e118fa6e7c8a"

测试情况

  • 失效的查询:针对第二个token的前缀搜索
    'channels like "c8a80b710e455460ae8b2399f5adfef5%"'
    
  • 正常的查询:中缀搜索
    'channels like "%c8a80b710e455460ae8b2399f5adfef5%"'
    
  • 正常的查询:第一个token的前缀搜索
    'channels like "48d302b1963841c39790fecf56b91ddc%"'
    

问题原因&解决方案

原因

Milvus中like "xxx%"语法是匹配整个字段的前缀,而非单个分词后的token前缀。你的channels字段按空白符分成了多个token,第二个token的前缀不是整个字段的开头,所以用like "xxx%"搜不到。而中缀%xxx%会匹配字段中任意位置的内容,所以能命中。

解决方法

要实现单个token的前缀搜索,需要使用Milvus的全文检索查询语法,配合INVERTED索引的分词能力:

  1. 调整查询语句:使用match查询并添加通配符*来匹配token前缀,示例:

    res = collection.query(
        expr='match(channels, "c8a80b710e455460ae8b2399f5adfef5*")',
        output_fields=["id", "channels"]
    )
    
  2. 优化索引配置:创建INVERTED索引时可以显式指定分词器(默认也是按空白符分词,显式配置更清晰),修改索引参数如下:

    {
      "field_name": "channels",
      "index_name": "channels_index",
      "index_type": "INVERTED",
      "params": {
        "analyzer": "simple"
      }
    }
    

    注意:修改索引需要重新创建集合或重建索引。

内容的提问来源于stack exchange,提问作者Qi Xiang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 10:52:32