You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Elasticsearch wildcard查询含数字字符串失效,如何配置命中目标文档?

问题描述

我有如下Elasticsearch文档:

{
  "_index" : "jobpipeline",
  "_type" : "_doc",
  "_id" : "xxxx",
  "_version" : 19,
  "_seq_no" : 417185,
  "_primary_term" : 18,
  "found" : true,
  "_source" : {
    "_boostScoreFactor" : 1.0,
    "type" : "jobpipeline",
    "owner" : [
      ""
    ],
    "title" : "xxxx",
    "watchTriggers" : [
      "uc4.abc",
      "uc4.abc.done"
    ],
    "touchTriggers" : [
      "uc4.abc.done"
    ]
  }
}

执行以下wildcard查询时无结果返回:

GET jobpipeline/_search
{
  "query": {
    "wildcard": {
      "touchTriggers": "*uc4.abc.done*"
    }
  }
}

但去掉uc4.后,使用以下查询可以命中该文档:

GET dd-jobpipeline/_search
{
  "query": {
    "wildcard": {
      "touchTriggers": "*abc.done*"
    }
  }
}

touchTriggers字段的映射如下:

"touchTriggers" : {
  "type" : "text",
  "fields" : {
    "keyword" : {
      "type" : "keyword",
      "ignore_above" : 256
    }
  }
}

JobPipeline索引的设置如下:

{
  "jobpipeline" : {
    "settings" : {
      "index" : {
        "routing" : {
          "allocation" : {
            "total_shards_per_node" : "2"
          }
        },
        "number_of_shards" : "1",
        "provided_name" : "jobpipeline",
        "creation_date" : "1675246845838",
        "unassigned" : {
          "node_left" : {
            "delayed_timeout" : "15m"
          }
        },
        "analysis" : {
          "filter" : {
            "discovery_stemmer" : {
              "type" : "stemmer",
              "language" : "porter2"
            },
            "discovery_stemmer_possessive" : {
              "type" : "stemmer",
              "language" : "minimal_english"
            },
            "discovery_synonym_abbr" : {
              "type" : "synonym_graph",
              "synonyms" : [
                "BN => browse node"
              ]
            }
          },
          "char_filter" : {
            "discovery_charfilter" : {
              "type" : "mapping",
              "mappings" : [
                ".=>|",
                "_=>|",
                "# => number",
                "% => percentage",
                "& => and",
                "+ => browse node"
              ]
            }
          },
          "normalizer" : {
            "lowercase" : {
              "filter" : [
                "lowercase"
              ],
              "type" : "custom",
              "char_filter" : [ ]
            }
          },
          "analyzer" : {
            "discovery_index_analyzer" : {
              "filter" : [
                "lowercase",
                "discovery_synonym_abbr"
              ],
              "char_filter" : [
                "discovery_charfilter"
              ],
              "type" : "custom",
              "tokenizer" : "standard"
            }
          }
        },
        "number_of_replicas" : "2",
        "uuid" : "FtylpFoeRFSG6QVBVdoYOg",
        "version" : {
          "created" : "135238227"
        }
      }
    }
  }
}

请问如何配置才能让*uc4.abc.done*的wildcard查询命中该文档?是否需要修改索引设置或映射?


原因分析

问题根源在索引的字符过滤器上。定义的discovery_charfilter会将所有.替换为|,导致uc4.abc.done在索引时被处理成uc4|abc|done。而wildcard查询基于分词后的词项匹配,查询时默认不会应用该字符过滤器,原查询中的.无法匹配索引中存储的|,因此无结果返回。

解决方案

有两种可行处理方式:

方式一:使用keyword子字段查询

touchTriggers的keyword子字段存储原始未分词字符串,直接针对该子字段执行wildcard查询即可,无需修改任何配置:

GET jobpipeline/_search
{
  "query": {
    "wildcard": {
      "touchTriggers.keyword": "*uc4.abc.done*"
    }
  }
}

方式二:调整字符过滤器或字段分析器

若必须使用text字段查询,需修改索引配置:

  1. 移除discovery_charfilter中.=>|的映射规则,避免.被替换;
  2. 或为touchTriggers单独指定不包含该字符过滤器的分析器,例如标准分析器:
PUT jobpipeline/_mapping
{
  "properties": {
    "touchTriggers": {
      "type": "text",
      "analyzer": "standard",
      "fields": {
        "keyword": {
          "type": "keyword",
          "ignore_above": 256
        }
      }
    }
  }
}

注意:修改映射后需要重新索引所有数据才能生效。


内容的提问来源于stack exchange,提问作者Jacky Guo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 02:36:06