Elasticsearch动态字段对象数组按key、value查询的实现与优化
你这个场景不需要自定义分词器,有两个性能远优于wildcard查询的成熟方案,具体如下:
方案一:nested类型预拆分(性能最优)
适用于查询频率高、对响应速度要求高的场景。
配置步骤
- 定义索引mapping,将
entities字段设为nested类型,内置key、value两个keyword字段:
{ "mappings": { "properties": { "entities": { "type": "nested", "properties": { "key": {"type": "keyword"}, "value": {"type": "keyword"} } } } } }
- 新增ingest pipeline,自动将原动态对象格式转成key-value结构:
{ "processors": [ { "script": { "source": """ def newEntities = []; for (entry in ctx.entities) { for (e in entry.entrySet()) { newEntities.add(["key": e.getKey(), "value": e.getValue()]); } } ctx.entities = newEntities; """ } } ] }
写入数据时指定该pipeline即可自动完成格式转换,无需修改上层写入逻辑。
三类查询实现
所有查询都走倒排索引,可达到毫秒级响应:
- 查询包含key为
propertyA的文档
{ "query": { "nested": { "path": "entities", "query": {"term": {"entities.key": "propertyA"}} } } }
- 查询包含value为
zxc的文档
{ "query": { "nested": { "path": "entities", "query": {"term": {"entities.value": "zxc"}} } } }
- 查询
propertyA值为qwe的文档
{ "query": { "nested": { "path": "entities", "query": { "bool": { "must": [ {"term": {"entities.key": "propertyA"}}, {"term": {"entities.value": "qwe"}} ] } } } } }
方案二:flattened类型(配置最简)
适用于不想修改写入逻辑、查询频率中等的场景。
配置步骤
直接把entities字段设为flattened类型即可,无需改动原有数据结构:
{ "mappings": { "properties": { "entities": {"type": "flattened"} } } }
三类查询实现
- 查询包含key为
propertyA的文档
{"query": {"exists": {"field": "entities.propertyA"}}}
- 查询包含value为
zxc的文档(需ES7.17及以上版本支持)
{"query": {"term": {"entities._extract_value": "zxc"}}}
- 查询
propertyA值为qwe的文档
{"query": {"term": {"entities.propertyA": "qwe"}}}
原方案性能差的原因
之前用拼接字段加wildcard的方案,因为wildcard查询无法完全利用倒排索引,通配符越多扫描的词条范围越大,性能自然差,上述两个方案都是直接利用ES原生索引结构,性能提升非常明显。
内容的提问来源于stack exchange,提问作者Егор Лебедев
相关产品推荐
相关产品推荐

