如何自定义Elasticsearch分词器以保留rtu格式编码不拆分?
Elasticsearch自定义分词器解决rtu开头编码被拆分问题
方案1:配置受保护模式的Word Delimiter Graph过滤器
默认分词器通过word_delimiter_graph_filter拆分短横线,我们可以自定义该过滤器,让它跳过以rtu开头的编码拆分:
创建带自定义分析器的索引
PUT /your_business_index { "settings": { "analysis": { "filter": { "rtu_protected_delimiter": { "type": "word_delimiter_graph", "preserve_original": true, "protected_patterns": [ "^rtu-\\w+(-\\w+)*$" ] } }, "analyzer": { "rtu_code_analyzer": { "tokenizer": "standard", "filter": [ "lowercase", "rtu_protected_delimiter" ] } } } }, "mappings": { "properties": { "device_code": { "type": "text", "analyzer": "rtu_code_analyzer", "search_analyzer": "rtu_code_analyzer" } } } }
protected_patterns用正则匹配所有rtu-xxx-xxx格式的编码,这类编码会被完整保留为单个token。preserve_original确保原始编码被存入索引,避免遗漏匹配可能。
方案2:使用Pattern分词器精准捕获编码
如果字段以这类编码为主,可直接用pattern分词器匹配完整编码:
PUT /your_business_index { "settings": { "analysis": { "analyzer": { "rtu_pattern_analyzer": { "type": "pattern", "pattern": "(rtu-\\w+(-\\w+)*)|(\\w+)", "group": 0 } } } }, "mappings": { "properties": { "device_code": { "type": "text", "analyzer": "rtu_pattern_analyzer" } } } }
- 正则优先匹配完整的rtu编码,再处理其他普通文本,确保编码不被拆分。
验证分词结果
用_analyze API测试效果:
POST /your_business_index/_analyze { "analyzer": "rtu_code_analyzer", "text": "rtu-2004-t89 office-301" }
返回结果中rtu-2004-t89会是单个token,而office-301仍会被正常拆分(如果不需要拆分其他短横线内容,可调整正则或过滤器配置)。
注意事项
- 已存在的索引无法修改分词配置,需重建索引并重新导入数据。
- 搜索时必须使用相同的分词器,避免搜索词被拆分导致匹配失败。
内容的提问来源于stack exchange,提问作者Klick
相关产品推荐
相关产品推荐

