Elasticsearch 8+ Painless脚本过滤问题:查询无法访问params._source
Elasticsearch嵌套数组过滤:保持字段关联匹配的解决方案
问题背景
需要过滤出some_availabilities.default.mark数组中存在满足以下条件的对象的文档:
online <= 指定timestampoffline >= 指定timestampsomedata1: truesomedata2: false
当前使用脚本过滤时,由于ES将数组字段拆分为独立存储的优化结构,无法保证不同字段数组的索引对应原始对象的顺序,导致条件判断出错。且新版本ES过滤脚本无法直接访问params._source,尝试设置nested类型未达预期。
映射配置
{ "mappings": { "properties": { "some_availabilities": { "properties": { "default": { "properties": { "mark": { "properties": { "somedata1": { "type": "boolean" }, "somedata2": { "type": "boolean" }, "offline": { "type": "long" }, "online": { "type": "long" } } }, "lily": { "type": "nested", "properties": { "somedata1": { "type": "boolean" }, "somedata2": { "type": "boolean" }, "offline": { "type": "long" }, "online": { "type": "long" } } } } } } } } } }
数据示例
{ "some_availabilities": { "default": { "lily": [], "mark": [ { "online": 1645480000, "offline": 1645481000, "somedata1": false, "somedata2": false }, { "online": 1645482000, "offline": 1645483000, "somedata1": true, "somedata2": false }, { "online": 1645484000, "offline": 1645485000, "somedata1": false, "somedata2": true }, { "online": 1645486000, "offline": 1645487000, "somedata1": true, "somedata2": true } ] } } }
当前尝试的查询(存在问题)
"query": { "bool": { "filter": { "script": { "script": { "source": "def result=false;long timestamp=params['timestamp'];def baseKey='some_availabilities.default.mark'; if(doc.containsKey(baseKey+ '.online')){def ctxAvailOnline = doc[baseKey + '.online'];for(int n=0;n<ctxAvailOnline.size();n++){long online;long offline;boolean field1;boolean field2;if(doc[baseKey+'.online'].size()>0){online=doc[baseKey+'.online'][n];offline=doc[baseKey+'.offline'][n];field1=doc[baseKey+'.somedata1'][n];field2=doc[baseKey+'.somedata2'][n];if(online<=timestamp&&offline>=timestamp&&field1&&!field2){result=true; return result}}}} ", "lang": "painless", "params": { "timestamp": 1645482000 } } } } } }
解决方案
1. 正确配置Nested类型(核心修复)
之前设置nested类型未生效,是因为**mark字段需要设置为nested类型**,而非上层的default。修改映射如下:
{ "mappings": { "properties": { "some_availabilities": { "properties": { "default": { "properties": { "mark": { "type": "nested", // 这里改为nested类型 "properties": { "somedata1": { "type": "boolean" }, "somedata2": { "type": "boolean" }, "offline": { "type": "long" }, "online": { "type": "long" } } }, "lily": { "type": "nested", "properties": { "somedata1": { "type": "boolean" }, "somedata2": { "type": "boolean" }, "offline": { "type": "long" }, "online": { "type": "long" } } } } } } } } } }
2. 使用Nested Query实现过滤
配置nested类型后,使用nested查询可以精准匹配数组内的独立对象,自动保证字段的关联关系,无需脚本即可实现需求:
{ "query": { "bool": { "filter": [ { "nested": { "path": "some_availabilities.default.mark", "query": { "bool": { "filter": [ { "range": { "some_availabilities.default.mark.online": { "lte": 1645482000 } } }, { "range": { "some_availabilities.default.mark.offline": { "gte": 1645482000 } } }, { "term": { "some_availabilities.default.mark.somedata1": true } }, { "term": { "some_availabilities.default.mark.somedata2": false } } ] } } } } ] } } }
3. 若必须使用脚本(不推荐,性能较差)
如果因特殊需求必须用脚本,可通过修改ES配置启用_source访问(会影响查询性能):
- 修改
elasticsearch.yml添加:script.field: _source: true - 重启ES后,脚本中可通过
params._source访问原始数据,示例脚本:
def markList = params._source.some_availabilities.default.mark; for (def item : markList) { if (item.online <= params.timestamp && item.offline >= params.timestamp && item.somedata1 && !item.somedata2) { return true; } } return false;
内容的提问来源于stack exchange,提问作者HeSh
相关产品推荐
相关产品推荐

