You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用Elasticsearch Ingest Pipeline实现多值字段的Enrich操作?

问题:为多值字段批量执行Enrich操作并返回数组结果

场景说明

现有Elasticsearch文档包含多值字段http.rule.id,示例结构:

{ "http.rule.id": ["b41912851a064912b2a589f3a21d0c57", "82045c5fd30045d893272fd8b74e93d6"] }

存在Enrich索引enrich-content,数据示例:

{
  "_index": "enrich-content",
  "_id": "b41912851a064912b2a589f3a21d0c57",
  "_source": {
    "description": "description1",
    "name": "name1",
    "location": "location1",
    "id": "b41912851a064912b2a589f3a21d0c57"
  }
},
{
  "_index": "enrich-content",
  "_id": "82045c5fd30045d893272fd8b74e93d6",
  "_source": {
    "description": "description2",
    "name": "name2",
    "location": "location2",
    "id": "82045c5fd30045d893272fd8b74e93d6"
  }
}

之前配置的两个Pipeline均仅能处理数组最后一个ID,无法得到目标结果http.description: ["description1", "description2"]。

失败原因分析

  • Pipeline 1:Enrich处理器默认max_matches=1,即使字段为数组,仅返回最后一个匹配结果,直接覆盖目标字段。
  • Pipeline 2:Foreach循环中每次执行Enrich都会覆盖http.description,最终仅保留最后一次的处理结果。

解决方案

方式一:利用Enrich处理器的max_matches参数(高效推荐)

通过设置max_matches为足够大的值,一次性获取所有匹配结果,再提取目标字段:

{
  "processors": [
    {
      "enrich": {
        "field": "http.rule.id",
        "policy_name": "policy_enrich",
        "target_field": "temp.enrich_results",
        "ignore_missing": true,
        "ignore_failure": true,
        "max_matches": 100  // 设置为大于业务场景中数组的最大长度
      }
    },
    {
      "script": {
        "source": """
          ctx.http.description = ctx.temp.enrich_results.stream().map(item -> item.description).collect(Collectors.toList());
          // 清理临时字段
          ctx.temp.remove("enrich_results");
          if (ctx.temp.isEmpty()) {
            ctx.remove("temp");
          }
        """,
        "ignore_failure": true
      }
    }
  ]
}

方式二:Foreach循环+Append处理器(逐个处理)

先初始化目标数组,循环处理每个ID后将结果追加到数组:

{
  "processors": [
    // 初始化目标数组为空
    {
      "set": {
        "field": "http.description",
        "value": [],
        "ignore_failure": true
      }
    },
    {
      "foreach": {
        "field": "http.rule.id",
        "processor": {
          "enrich": {
            "field": "_ingest._value",
            "policy_name": "policy_enrich",
            "target_field": "_temp",
            "ignore_missing": true,
            "ignore_failure": true
          }
        },
        "ignore_failure": true
      }
    },
    // 将临时字段中的description追加到目标数组
    {
      "foreach": {
        "field": "_temp",
        "processor": {
          "append": {
            "field": "http.description",
            "value": "{{_ingest._value.description}}"
          }
        },
        "ignore_failure": true
      }
    },
    // 删除临时字段
    {
      "remove": {
        "field": "_temp",
        "ignore_failure": true
      }
    }
  ]
}

注意事项

  • 确保Enrich策略policy_enrich的匹配字段为id,源字段包含description。
  • 方式一的max_matches值需根据业务中http.rule.id的最大数组长度调整,避免遗漏数据。

内容的提问来源于stack exchange,提问作者Dmitry

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 14:53:15