如何利用Elasticsearch Ingest Pipeline实现多值字段的Enrich操作?
问题:为多值字段批量执行Enrich操作并返回数组结果
场景说明
现有Elasticsearch文档包含多值字段http.rule.id,示例结构:
{ "http.rule.id": ["b41912851a064912b2a589f3a21d0c57", "82045c5fd30045d893272fd8b74e93d6"] }
存在Enrich索引enrich-content,数据示例:
{ "_index": "enrich-content", "_id": "b41912851a064912b2a589f3a21d0c57", "_source": { "description": "description1", "name": "name1", "location": "location1", "id": "b41912851a064912b2a589f3a21d0c57" } }, { "_index": "enrich-content", "_id": "82045c5fd30045d893272fd8b74e93d6", "_source": { "description": "description2", "name": "name2", "location": "location2", "id": "82045c5fd30045d893272fd8b74e93d6" } }
之前配置的两个Pipeline均仅能处理数组最后一个ID,无法得到目标结果http.description: ["description1", "description2"]。
失败原因分析
- Pipeline 1:Enrich处理器默认
max_matches=1,即使字段为数组,仅返回最后一个匹配结果,直接覆盖目标字段。 - Pipeline 2:Foreach循环中每次执行Enrich都会覆盖
http.description,最终仅保留最后一次的处理结果。
解决方案
方式一:利用Enrich处理器的max_matches参数(高效推荐)
通过设置max_matches为足够大的值,一次性获取所有匹配结果,再提取目标字段:
{ "processors": [ { "enrich": { "field": "http.rule.id", "policy_name": "policy_enrich", "target_field": "temp.enrich_results", "ignore_missing": true, "ignore_failure": true, "max_matches": 100 // 设置为大于业务场景中数组的最大长度 } }, { "script": { "source": """ ctx.http.description = ctx.temp.enrich_results.stream().map(item -> item.description).collect(Collectors.toList()); // 清理临时字段 ctx.temp.remove("enrich_results"); if (ctx.temp.isEmpty()) { ctx.remove("temp"); } """, "ignore_failure": true } } ] }
方式二:Foreach循环+Append处理器(逐个处理)
先初始化目标数组,循环处理每个ID后将结果追加到数组:
{ "processors": [ // 初始化目标数组为空 { "set": { "field": "http.description", "value": [], "ignore_failure": true } }, { "foreach": { "field": "http.rule.id", "processor": { "enrich": { "field": "_ingest._value", "policy_name": "policy_enrich", "target_field": "_temp", "ignore_missing": true, "ignore_failure": true } }, "ignore_failure": true } }, // 将临时字段中的description追加到目标数组 { "foreach": { "field": "_temp", "processor": { "append": { "field": "http.description", "value": "{{_ingest._value.description}}" } }, "ignore_failure": true } }, // 删除临时字段 { "remove": { "field": "_temp", "ignore_failure": true } } ] }
注意事项
- 确保Enrich策略
policy_enrich的匹配字段为id,源字段包含description。 - 方式一的
max_matches值需根据业务中http.rule.id的最大数组长度调整,避免遗漏数据。
内容的提问来源于stack exchange,提问作者Dmitry
相关产品推荐
相关产品推荐

