You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Elasticsearch分析器在API测试正常但搜索查询无预期结果

Elasticsearch查询时无法返回分析器分词结果的问题解决

问题背景

我在索引的settings和mapping中配置了自定义分析器regex_analyzer,配置内容如下:

{
    "settings": {
        "index": {
            "analysis": {
                "analyzer": {
                    "synonym_analyzer": {
                        "tokenizer": "standard",
                        "filter": [
                            "lowercase"
                        ]
                    },
                    "regex_analyzer": {
                        "tokenizer": "regex_tokenizer",
                        "filter": [
                            "lowercase"
                        ]
                    }
                },
                "tokenizer": {
                    "regex_tokenizer": {
                        "type": "pattern",
                        "pattern": "((\\b|\\s|\\.|,)[a-z](\\b|\\s |\\.|,)){3,}",
                        "group": 0
                    }
                }
            }
        }
    },
    "mappings": {
        "properties": {
            "transcript_data": {
                "properties": {
                    "transcript": {
                        "type": "text",
                        "fields": {
                            "keyword": {
                                "type": "keyword"
                            },
                            "regex": {
                                "type": "text",
                                "analyzer": "regex_analyzer",
                                "search_analyzer": "regex_analyzer"
                            }
                        }
                    }
                }
            }
        }
    }
}

直接调用_analyze API测试该分析器时,能得到符合预期的分词结果:

POST myIndex/_analyze
{
  "analyzer": "regex_analyzer",
  "text": " this article is talking about l a z r and b k k t ...."
}

响应结果:

{
  "tokens" : [
    {
      "token" : " b k k t",
      "start_offset" : 7971,
      "end_offset" : 7979,
      "type" : "word",
      "position" : 0
    },
    {
      "token" : " l a z r",
      "start_offset" : 8350,
      "end_offset" : 8358,
      "type" : "word",
      "position" : 1
    }
  ]
}

但使用match_all查询并指定返回transcript_data.transcript.regex字段时,返回的是完整的原始文本,而非分词后的数组:

GET myIndex/_search
{
  "query": {
   "match_all": {

   }
  },
  "fields": [
    "transcript_data.transcript.regex"
  ]
}

响应结果中fields部分:

"fields" : {
  "transcript_data.transcript.regex" : [
    " this article is talking about l a z r and b k k t ...."
  ]
}

我原本期望这个字段返回与_analyze一致的分词结果。

原因分析

Elasticsearch的fields参数返回的是字段的原始值(仅经过索引时的字符过滤器处理,不会拆解成分词),倒排索引中的分词tokens是用于搜索匹配的内部数据,默认不会通过_search的fields直接返回。

解决方案

方案1:使用Term Vectors API获取分词结果

Term Vectors API可以直接返回指定文档中某个字段的分词信息,包括token、位置、偏移量等,完全匹配_analyze的输出格式。调用方式如下:

GET myIndex/_termvectors/46?fields=transcript_data.transcript.regex

其中46是目标文档的_id,响应结果会包含该字段经过regex_analyzer处理后的所有分词tokens。

方案2:在索引时存储分词结果(可选)

如果需要在搜索时直接返回分词结果,可以通过Ingest Pipeline在文档写入前完成分词,将结果存入一个新字段:

  1. 创建分词处理的pipeline:
PUT _ingest/pipeline/extract_regex_tokens
{
  "processors": [
    {
      "script": {
        "source": """
          def analyzer = ctx._index + '_' + ctx._type + '#regex_analyzer';
          def tokens = analyze(analyzer, ctx.transcript_data.transcript);
          ctx.transcript_data.regex_tokens = tokens.stream().map(token -> token.token).collect(Collectors.toList());
        """
      }
    }
  ]
}
  1. 写入文档时指定pipeline:
PUT myIndex/_doc/46?pipeline=extract_regex_tokens
{
  "doc_type" : "post",
  "transcript_data" : {
    "transcript" : "this article is talking about l a z r and b k k t ...."
  },
  "join_field" : {
    "name" : "video",
    "parent" : "anonymouse"
  }
}

之后搜索时就可以直接返回transcript_data.regex_tokens字段,得到分词后的数组。

内容的提问来源于stack exchange,提问作者Exorcismus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 18:54:27