Elasticsearch如何仅筛选一级字符串/数字类型的原始属性?
解决方案:用Ingest Pipeline动态筛选一级原始类型字段
这问题我之前处理过类似场景,动态字段多的时候,_source的includes/excludes确实不好用,这种情况下Ingest Pipeline的Script Processor是最优解,能帮你动态判断一级字段的类型,只保留字符串和数字类型的字段。
具体步骤:
创建自定义Ingest Pipeline
我们用Painless脚本遍历文档的所有一级字段,只保留值为字符串(String)或数字(Number,包含int、long、float等所有数字类型)的字段。执行以下命令创建pipeline:PUT _ingest/pipeline/keep_primitive_top_level_fields { "processors": [ { "script": { "source": """ def filteredCtx = [:]; // 遍历所有一级字段 for (Map.Entry entry in ctx.entrySet()) { def fieldValue = entry.getValue(); // 判断是否为字符串或数字类型 if (fieldValue instanceof String || fieldValue instanceof Number) { filteredCtx[entry.getKey()] = fieldValue; } } // 替换原文档为筛选后的内容 ctx = filteredCtx; """ } } ] }索引文档时指定该Pipeline
索引你的JSON文档时,通过pipeline参数指定我们刚创建的pipeline,这样文档在写入Elasticsearch前会自动被处理:PUT your-index-name/_doc/1?pipeline=keep_primitive_top_level_fields { "id":"1234", "ad":{ "id": "50419957", "priority": 1 }, "div":["adb", "dfg"], "startDate": 1533193200000, "endDate": 1534489199000 }验证结果
执行查询命令获取文档:GET your-index-name/_doc/1返回的结果会只包含
id、startDate、endDate三个字段,嵌套对象ad和数组div都会被过滤掉,完全符合你的需求。
额外扩展
如果之后需要排除某些特定的一级字段(即使它是字符串/数字类型),只需在脚本里加个判断即可,比如要排除名为ignore_me的字段:
if ((fieldValue instanceof String || fieldValue instanceof Number) && !entry.getKey().equals("ignore_me")) { filteredCtx[entry.getKey()] = fieldValue; }
内容的提问来源于stack exchange,提问作者Renjith
相关产品推荐
相关产品推荐

