You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Elasticsearch如何仅筛选一级字符串/数字类型的原始属性?

解决方案:用Ingest Pipeline动态筛选一级原始类型字段

这问题我之前处理过类似场景,动态字段多的时候,_source的includes/excludes确实不好用,这种情况下Ingest Pipeline的Script Processor是最优解,能帮你动态判断一级字段的类型,只保留字符串和数字类型的字段。

具体步骤:

  1. 创建自定义Ingest Pipeline
    我们用Painless脚本遍历文档的所有一级字段,只保留值为字符串(String)或数字(Number,包含int、long、float等所有数字类型)的字段。执行以下命令创建pipeline:

    PUT _ingest/pipeline/keep_primitive_top_level_fields
    {
      "processors": [
        {
          "script": {
            "source": """
              def filteredCtx = [:];
              // 遍历所有一级字段
              for (Map.Entry entry in ctx.entrySet()) {
                def fieldValue = entry.getValue();
                // 判断是否为字符串或数字类型
                if (fieldValue instanceof String || fieldValue instanceof Number) {
                  filteredCtx[entry.getKey()] = fieldValue;
                }
              }
              // 替换原文档为筛选后的内容
              ctx = filteredCtx;
            """
          }
        }
      ]
    }
    
  2. 索引文档时指定该Pipeline
    索引你的JSON文档时,通过pipeline参数指定我们刚创建的pipeline,这样文档在写入Elasticsearch前会自动被处理:

    PUT your-index-name/_doc/1?pipeline=keep_primitive_top_level_fields
    { 
      "id":"1234", 
      "ad":{ "id": "50419957", "priority": 1 }, 
      "div":["adb", "dfg"], 
      "startDate": 1533193200000, 
      "endDate": 1534489199000 
    }
    
  3. 验证结果
    执行查询命令获取文档:

    GET your-index-name/_doc/1
    

    返回的结果会只包含id、startDate、endDate三个字段,嵌套对象ad和数组div都会被过滤掉,完全符合你的需求。

额外扩展

如果之后需要排除某些特定的一级字段(即使它是字符串/数字类型),只需在脚本里加个判断即可,比如要排除名为ignore_me的字段:

if ((fieldValue instanceof String || fieldValue instanceof Number) && !entry.getKey().equals("ignore_me")) {
  filteredCtx[entry.getKey()] = fieldValue;
}

内容的提问来源于stack exchange,提问作者Renjith

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:11:02