NiFi中JOLTTransformJson处理器结构转换不符合预期求助
NiFi JOLTTransformJson 输出不符合预期问题排查
输入JSON
[ { "col_name": "time", "data_type": "timestamp", "is_nullable": true }, { "col_name": "otherData", "data_type": "string", "is_nullable": false } ]
当前使用的JOLT规格
[ { "operation": "shift", "spec": { "*": { "col_name": "name", "data_type": "type[0]", "is_nullable": { "true": "type[1]", "false": "type[1]" } } } }, { "operation": "default", "spec": { "*": { "type[1]": "notnull" } } } ]
预期输出
{ "type": "record", "name": "table_name", "fields": [ { "name": "time", "type": [ "timestamp", "null" ] }, { "name": "otherData", "type": [ "string", "notnull" ] } ] }
实际输出
{ "name": [ "time", "otherData" ], "type": [ [ "timestamp", "int" ], null ] }
问题分析与修正方案
原JOLT的核心问题是shift操作未为每个输入数组元素创建独立的对象结构,而是直接将同名字段的值合并到顶层数组中,导致结构完全不符合预期。需要调整shift操作的映射规则,为每个输入元素生成fields数组内的独立对象,同时处理is_nullable的分支逻辑,最后补充顶层的固定字段。
修正后的JOLT规格如下:
[ { "operation": "shift", "spec": { "*": { // 用&1引用外层数组索引,将每个元素映射到fields的对应位置 "col_name": "fields[&1].name", "data_type": "fields[&1].type[0]", "is_nullable": { "true": "fields[&1].type[1]", // 可空时type[1]设为null "false": "fields[&1].type[1]" // 不可空时type[1]设为notnull } } } }, { "operation": "modify-overwrite-beta", "spec": { "fields": { "*": { // 为is_nullable=true的元素设置type[1]为null "type[1]": ["=isNull", "null", "@(1,type[1])"], // 为is_nullable=false的元素设置type[1]为notnull(兜底) "type[1]": ["=isNull", "notnull", "@(1,type[1])"] } } } }, { "operation": "default", "spec": { // 添加顶层固定字段 "type": "record", "name": "table_name" } } ]
规则说明
- shift操作:通过
&1(外层数组的索引)将每个输入元素映射到fields数组的对应位置,确保每个字段都归属到独立的对象中。 - modify-overwrite-beta操作:处理
is_nullable的分支值,当is_nullable=true时将type[1]设为null,否则设为notnull,避免default操作的全局覆盖问题。 - default操作:添加顶层的
type: "record"和name: "table_name"固定字段,确保输出结构完整。
使用修正后的JOLT规格,即可得到符合预期的输出结果。
内容的提问来源于stack exchange,提问作者Ashu Tech
相关产品推荐
相关产品推荐

