You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Apache NiFi的ConvertRecord将含动态字段的JSON转CSV?

使用Apache NiFi的ConvertRecord处理动态字段JSON转CSV

嘿,这个需求用ConvertRecord完全能搞定,而且比替换或者JSON求值靠谱多了——我来给你一步步拆解,保证你能弄明白:

第一步:定义兼容所有动态字段的Record Schema(核心!)

因为你的JSON记录最多有11个字段,每条记录的字段数量不固定,所以我们需要先定义一个包含全部11个字段的Schema,并且把每个字段设为可选(允许缺失)。这里用Avro Schema举例,这也是NiFi里最常用的Schema格式:

{
  "type": "record",
  "name": "DynamicJsonToCsvSchema",
  "fields": [
    {"name": "field1", "type": ["null", "string"], "default": null},
    {"name": "field2", "type": ["null", "string"], "default": null},
    {"name": "field3", "type": ["null", "string"], "default": null},
    {"name": "field4", "type": ["null", "string"], "default": null},
    {"name": "field5", "type": ["null", "string"], "default": null},
    {"name": "field6", "type": ["null", "string"], "default": null},
    {"name": "field7", "type": ["null", "string"], "default": null},
    {"name": "field8", "type": ["null", "string"], "default": null},
    {"name": "field9", "type": ["null", "string"], "default": null},
    {"name": "field10", "type": ["null", "string"], "default": null},
    {"name": "field11", "type": ["null", "string"], "default": null}
  ]
}

解释下:每个字段用["null", "string"]的联合类型,再加上default: null,这样当某条记录没有这个字段时,NiFi会自动填充null,不会因为字段缺失报错。如果你的字段有其他类型(比如数字、布尔),把string换成对应的类型就行。

第二步:配置ConvertRecord处理器

现在来配置ConvertRecord,别担心,其实就是选对Reader和Writer,再绑定Schema:

  • 配置Record Reader(解析JSON)
    选择JsonReader(或者JsonTreeReader,两者都能用,JsonReader更轻量),然后设置:

    • Schema Access Strategy:选Use Schema Text,然后把上面的Avro Schema直接粘贴到Schema Text属性里。
    • Allow Extra Fields:设为true(防止万一有超出11个的字段,处理器不会报错,只是忽略)。
  • 配置Record Writer(生成CSV)
    选择CSVRecordSetWriter,这是NiFi专门用来生成标准CSV的Writer,关键配置:

    • Schema Access Strategy:选Use Schema from Record Reader,这样Writer会自动复用Reader的Schema来生成CSV的表头和列顺序。
    • Include Header Line:设为true,这样输出的CSV会有完整的表头(field1到field11)。
    • Null Value:设为空字符串"",这样缺失的字段在CSV里显示为空,而不是null,更符合CSV的使用习惯。
    • CSV Format:默认选RFC4180就行,这是通用的CSV标准格式。
  • 其他通用配置
    剩下的属性比如Batch Size、Character Set保持默认就好,除非你有特殊的批量处理或编码需求。

第三步:验证处理效果

举个实际的例子,假设你有三条动态字段的JSON:

{"field1":"text1", "field3":"text3", "field7":"text7"}
{"field1":"textA", "field2":"textB", "field4":"textD", "field5":"textE", "field6":"textF", "field8":"textH", "field9":"textI", "field10":"textJ", "field11":"textK"}
{"field2":"test2", "field5":"test5", "field9":"test9"}

经过ConvertRecord处理后,生成的CSV会是这样的:

field1,field2,field3,field4,field5,field6,field7,field8,field9,field10,field11
text1,,text3,,,,text7,,,,
textA,textB,,textD,textE,textF,,,,textH,textI,textJ,textK
,test2,,,test5,,,,test9,,

完美!缺失的字段自动留空,表头包含所有11个字段,完全满足你的动态字段需求。

小技巧

如果你不想手动写Schema,可以先用InferSchema处理器从你的样本JSON数据自动生成Schema,然后再手动给每个字段加上null类型和默认值,这样更高效。

内容的提问来源于stack exchange,提问作者meh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:52:48