You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Elasticsearch自定义分析器停用词过滤与文本存储异常问题

Elasticsearch自定义分析器疑问解答

操作过程

创建索引请求

PUT my_index
{
  "settings": {
    "index.number_of_shards": 1,
    "index.number_of_replicas": 0,
    "analysis": {
      "analyzer": {
        "my_custom_analyzer": {
          "type": "custom",
          "tokenizer": "standard",
          "max_token_length": 3,
          "filter": [
            "lowercase",
            "english_stop",
            "english_stemmer",
            "asciifolding"
          ]
        }
      },
      "filter": {
        "english_stemmer": {
          "type": "stemmer",
          "stopwords": "english"
        },
        "english_stop": {
          "type": "stop",
          "stopwords": "_english_"
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "Text": {
        "type": "text",
        "analyzer": "my_custom_analyzer"
      }
    }
  }
}

添加文档请求

POST my_index/_doc/1
{
  "Text": "a quick fox jumps over the lazy dog"
}

查询请求

GET my_index/_search
{
  "query": {
    "match": {
      "Text": {
        "query": "a the", 
        "operator": "and"
      }
    }
  }
}

查询响应

{
  "took" : 1,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 1,
      "relation" : "eq"
    },
    "max_score" : 0.5753642,
    "hits" : [
      {
        "_index" : "my_index",
        "_type" : "_doc",
        "_id" : "1",
        "_score" : 0.5753642,
        "_source" : {
          "Text" : "a quick fox jumps over the lazy dog"
        }
      }
    ]
  }
}

疑问与解答

疑问1:停用词查询为何仍能匹配结果?

原因是查询阶段会使用字段指定的分析器处理查询文本。你定义的my_custom_analyzer包含停用词过滤器english_stop,当查询文本"a the"被分析时,这两个停用词会被直接过滤掉,最终查询没有有效匹配词。此时match查询(operator: and)会退化为匹配所有文档,所以你能得到结果。

可以用_analyze接口验证分析器对查询文本的处理结果:

GET _analyze
{
  "analyzer": "my_custom_analyzer",
  "text": "a the"
}

返回结果会显示没有生成任何词项,这直接解释了查询的行为。

疑问2:Text字段为何仍存储原始文本?

Elasticsearch中text字段的_source是原始文档的备份,用于查询时返回给用户查看,它不会被分析器修改。真正用于检索的是索引中存储的经过分析后的词项(也就是你预期的去除停用词后的精简版本),这些词项仅用于内部匹配计算,不会直接返回。

可以用_termvectors接口确认索引中的词项:

GET my_index/_doc/1/_termvectors
{
  "fields": ["Text"]
}

返回结果里不会包含"a"和"the"这些停用词,证明分析器确实在索引阶段过滤了它们。


内容的提问来源于stack exchange,提问作者No1Lives4Ever

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 05:26:46