Painless中doc.size()/keySet()报UnsupportedOperationException及文档属性重映射问题
问题解答
a)为何无法访问doc的这些方法
这是因为脚本上下文的限制:你使用的filter脚本上下文里,doc对象是LeafDocLookup类型——它是OpenSearch为高性能过滤场景专门优化的实现,仅支持直接访问具体字段(比如doc['gpa'].value),不支持size()、keySet()这类遍历/统计操作。
原因在于,filter上下文的核心目标是快速过滤文档,LeafDocLookup采用延迟加载机制,只有显式访问某个字段时才会加载对应数据。如果允许遍历所有字段,会触发全字段加载,完全违背该上下文的性能优化设计。
只有在部分其他上下文(比如scripted_metric聚合的map阶段、update脚本)中,才能使用支持遍历的文档访问方式(比如ctx._source)。
b)无需显式列出所有属性完成重映射的方案
根据需求,推荐两种可行方案,选择哪种取决于你是要查询时动态转换还是索引时永久转换:
方案1:用scripted_metric聚合实现查询时动态转换
在scripted_metric的map阶段,可通过ctx._source访问文档原始字段,它支持遍历所有键值对,能动态识别并处理所有x.y.开头的字段。
示例查询脚本:
POST /testindex1/_search { "size": 0, "aggs": { "remapped_documents": { "scripted_metric": { "init_script": "state.docs = []", "map_script": """ // 构建props对象存储重映射后的字段 Map props = new HashMap(); // 遍历文档所有原始字段 for (Map.Entry entry in ctx._source.entrySet()) { String fieldName = entry.getKey(); if (fieldName.startsWith('x.y.')) { // 去掉前缀x.y.,得到新键名 String newFieldName = fieldName.substring(4); props[newFieldName] = entry.getValue(); } } // 组装最终的重映射文档 Map remappedDoc = new HashMap(); remappedDoc['props'] = props; // 保留其他需要的原始字段(可选) if (ctx._source.containsKey('grad')) { remappedDoc['grad'] = ctx._source['grad']; } if (ctx._source.containsKey('gpa')) { remappedDoc['gpa'] = ctx._source['gpa']; } state.docs.add(remappedDoc); """, "combine_script": "return state.docs", "reduce_script": "return states.flatten()" } } } }
方案2:用Ingest Pipeline实现索引时永久转换
如果希望文档在索引时就完成重映射(避免每次查询都处理),推荐使用Ingest Pipeline,这是更高效的方式:
- 创建处理管道:
PUT /_ingest/pipeline/remap_x_y_fields { "processors": [ { "script": { "source": """ Map props = new HashMap(); // 迭代文档字段,避免遍历过程中修改集合引发异常 Iterator<Map.Entry> fieldIterator = ctx.entrySet().iterator(); while (fieldIterator.hasNext()) { Map.Entry fieldEntry = fieldIterator.next(); String fieldName = fieldEntry.getKey(); if (fieldName.startsWith('x.y.')) { String newFieldName = fieldName.substring(4); props[newFieldName] = fieldEntry.getValue(); // 移除原始的x.y.开头字段(可选) fieldIterator.remove(); } } // 将props添加到文档中 ctx['props'] = props; """ } } ] }
- 索引文档时指定管道:
POST /testindex1/_doc?pipeline=remap_x_y_fields { "x.y.a": "testValueA", "x.y.b": 456, "grad": true, "gpa": 3.79 }
索引完成后,文档会直接包含props.a、props.b字段,后续查询无需额外处理。
内容的提问来源于stack exchange,提问作者Pete
相关产品推荐
相关产品推荐

