You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Elasticsearch的text类型是否有类似ignore_below的短词过滤配置?

解决方案

你可以通过Elasticsearch的自定义分析器搭配length令牌过滤器实现需求,该过滤器可以直接筛除长度不符合要求的分词结果。

1. 完整索引配置示例

需要同时修改索引的settings(定义自定义分析器)和mappings(给publication字段绑定分析器),参考配置如下:

PUT 你的索引名称
{
  "settings": {
    "analysis": {
      "filter": {
        "min_length_filter": {
          "type": "length",
          "min": 3,
          "max": 255
        }
      },
      "analyzer": {
        "ignore_short_word_analyzer": {
          "tokenizer": "standard",
          "filter": [
            "lowercase",
            "min_length_filter"
          ]
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "publication": {
        "type": "text",
        "analyzer": "ignore_short_word_analyzer",
        "fields": {
          "keyword": {
            "type": "keyword"
          }
        }
      },                
      "published": {
        "type": "date",
        "fields": {
          "keyword": { 
            "type": "keyword"
          }
        }
      },
      "id": {
        "type": "text"
      }
    }
  }
}

2. 效果验证

配置完成后可以用分析接口测试分词效果:

POST 你的索引名称/_analyze
{
  "analyzer": "ignore_short_word_analyzer",
  "text": "A car with an add"
}

返回的分词结果就是你需要的["car", "with", "add"]。

3. 已有索引处理说明

如果索引已经创建完成,无法直接修改分析器配置,需要按以下步骤操作:

  • 关闭索引:POST /你的索引名称/_close
  • 更新索引settings添加上述自定义分析器配置
  • 重新打开索引:POST /你的索引名称/_open
  • 对存量数据执行重新索引,新的分词规则才会生效

内容的提问来源于stack exchange,提问作者akrelj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 04:27:01