Elasticsearch中icu_collation_keyword字段排序异常问题求助
Elasticsearch icu_collation_keyword 排序异常问题排查与解决建议
问题描述
使用Elasticsearch的icu_collation_keyword字段类型对文档排序时出现两个异常:
- 排序顺序不符合预期的大小写敏感排序规则;
- 响应结果的
sort字段显示乱码字符。
复现步骤
创建测试索引
PUT /test-index { "mappings": { "properties": { "id422": { "type": "text", "fields": { "collated": { "type": "icu_collation_keyword", "strength": "tertiary", "case_level": true } } } } } }
写入测试文档
POST /test-index/_doc/1 { "id422": "0a11" } POST /test-index/_doc/2 { "id422": "0A11" } POST /test-index/_doc/3 { "id422": "0b11" } POST /test-index/_doc/4 { "id422": "0B11" } POST /test-index/_doc/5 { "id422": "0c11" } POST /test-index/_doc/6 { "id422": "0C11" }
执行排序搜索
GET /test-index/_search { "sort": [ { "id422.collated": { "order": "asc" } } ], "_source": ["id422"] }
预期与实际结果
预期排序顺序
0A11 0B11 0C11 0a11 0b11 0c11
实际排序顺序
0a11 0A11 0b11 0B11 0c11 0C11
同时响应的sort字段出现乱码字符(如কՅ‡ࡀ)。
补充信息
- Elasticsearch版本:8.15.3
- Kibana版本:8.15.3
- ICU Analysis插件版本:8.15.3
解决思路与建议
1. 调整排序参数实现预期大小写顺序
当前映射仅设置了strength: tertiary(区分大小写、重音)和case_level: true(启用大小写级别排序),但未指定大小写的优先顺序,默认规则是小写字母排在大写字母之前。需要添加case_first: "upper"参数,让大写字母优先排序。
修改后的索引映射如下:
PUT /test-index { "mappings": { "properties": { "id422": { "type": "text", "fields": { "collated": { "type": "icu_collation_keyword", "strength": "tertiary", "case_level": true, "case_first": "upper" } } } } } }
修改映射后,需要重新创建索引并写入测试文档,再执行排序搜索即可得到预期的顺序。
2. 关于sort字段的乱码问题
icu_collation_keyword字段存储的是二进制排序键,而非原始字符串,这些乱码字符是排序键的二进制值转义后的显示,属于正常现象,不影响排序逻辑,无需关注。如果需要在响应中显示原始排序字段值,可以在搜索请求中指定sort的unmapped_type或直接使用原始字段配合排序参数,但icu_collation_keyword的优势就是预计算排序键提升性能,建议保留当前字段类型,忽略sort字段的乱码显示。
内容的提问来源于stack exchange,提问作者user28262719
相关产品推荐
相关产品推荐

