Elasticsearch 7.9.1中ngram tokenizer的max_ngram未生效求助
解决Elasticsearch 7.9.1中Ngram Tokenizer不生成长Ngram的问题
问题原因
你的索引配置中,ngram_tokenizer的min_ngram和max_ngram参数使用了字符串类型(带引号的"2"和"40"),但Elasticsearch要求这两个参数必须是整数类型。这导致配置未被正确解析,Elasticsearch使用了Ngram Tokenizer的默认值:min_ngram=1、max_ngram=2,所以只会生成1-2字符的token。
解决步骤
由于Elasticsearch创建索引后无法修改分析器配置,需要删除现有索引并重新创建:
- 删除旧索引
curl -X DELETE 'http://elasticsearch_test:9200/test'
- 创建新索引(使用正确的整数参数)
{ "settings": { "index": { "max_ngram_diff": 40, "number_of_shards": 1, "number_of_replicas": 1, "analysis": { "analyzer": { "ngram_analyzer": { "filter": [ "lowercase", "trim" ], "type": "custom", "tokenizer": "ngram_tokenizer" }, "search_analyzer": { "filter": [ "lowercase", "trim" ], "type": "custom", "tokenizer": "keyword" } }, "tokenizer": { "ngram_tokenizer": { "type": "ngram", "max_ngram": 40, "min_ngram": 2 } } } } } }
执行创建命令:
curl -X PUT 'http://elasticsearch_test:9200/test' \ --header 'Content-Type: application/json' \ --data '@上述配置文件路径'
- 重新测试分析器
使用你原来的测试命令:
curl --location --request GET 'http://elasticsearch_test:9200/test/_analyze?pretty=true' \ --header 'Content-Type: application/json' \ --data '{ "analyzer": "ngram_analyzer", "text": "cod_1" }'
预期结果
此时会生成所有长度从2到5(输入文本总长度为5)的ngram,示例返回如下:
{ "tokens" : [ { "token" : "co", "start_offset" : 0, "end_offset" : 2, "type" : "word", "position" : 0 }, { "token" : "cod", "start_offset" : 0, "end_offset" : 3, "type" : "word", "position" : 1 }, { "token" : "cod_", "start_offset" : 0, "end_offset" : 4, "type" : "word", "position" : 2 }, { "token" : "cod_1", "start_offset" : 0, "end_offset" : 5, "type" : "word", "position" : 3 }, { "token" : "od", "start_offset" : 1, "end_offset" : 3, "type" : "word", "position" : 4 }, { "token" : "od_", "start_offset" : 1, "end_offset" : 4, "type" : "word", "position" : 5 }, { "token" : "od_1", "start_offset" : 1, "end_offset" : 5, "type" : "word", "position" : 6 }, { "token" : "d_", "start_offset" : 2, "end_offset" : 4, "type" : "word", "position" : 7 }, { "token" : "d_1", "start_offset" : 2, "end_offset" : 5, "type" : "word", "position" : 8 }, { "token" : "_1", "start_offset" : 3, "end_offset" : 5, "type" : "word", "position" : 9 } ] }
内容的提问来源于stack exchange,提问作者Olego
相关产品推荐
相关产品推荐

