在Druid中摄入Base64编码Avro消息遇EOFException的解决问询
问题
在Druid中摄入Base64编码的Avro消息时,遇到错误:Avro's unnecessary EOFException, detail: https://issues.apache.org/jira/browse/AVRO-813
查看Druid的InlineSchemaAvroBytesDecoder.java代码第88行后,发现该解码器未对消息做Base64解码处理。请问是否遗漏了相关配置?如何配置Druid以解析Base64编码的Avro消息?
当前使用的输入格式配置Spec如下:
"inputFormat": { "type": "avro_stream", "avroBytesDecoder": { "type": "schema_inline", "schema": { "namespace": "org.apache.druid.data", "name": "User", "type": "record", "fields": [ { "name": "id", "type": "string" }, { "name": "price", "type": "int" } ] } }, "flattenSpec": { "useFieldDiscovery": true, "fields": [ { "type": "path", "name": "someRecord_subInt", "expr": "$.someRecord.subInt" } ] }, "binaryAsString": false }
解决方法
Druid默认的Avro解码器不支持直接处理Base64编码的消息,需要在摄入流程中添加Base64解码步骤,具体配置调整如下:
针对带JSON结构的输入(Base64字段嵌套在JSON中)
切换输入格式为json,通过flattenSpec的transform功能对Base64字段解码,再用Avro解析二进制数据:"inputFormat": { "type": "json", "flattenSpec": { "useFieldDiscovery": true, "fields": [ { "type": "path", "name": "avroBinary", "expr": "$.yourBase64FieldName", "transform": { "type": "base64Decode" } } ] }, "avroBytesDecoder": { "type": "schema_inline", "schema": { "namespace": "org.apache.druid.data", "name": "User", "type": "record", "fields": [ { "name": "id", "type": "string" }, { "name": "price", "type": "int" } ] } }, "binaryAsString": false }注意替换
yourBase64FieldName为实际存储Base64编码Avro的字段名。针对纯Base64字符串输入(无JSON结构)
使用text输入格式配合表达式转换,先解码Base64再解析Avro:"inputFormat": { "type": "text", "transformSpec": { "transforms": [ { "type": "expression", "name": "avroBinary", "expression": "base64Decode(input)" } ] }, "avroBytesDecoder": { "type": "schema_inline", "schema": { "namespace": "org.apache.druid.data", "name": "User", "type": "record", "fields": [ { "name": "id", "type": "string" }, { "name": "price", "type": "int" } ] } } }
内容的提问来源于stack exchange,提问作者Rahul Kumar Gond
相关产品推荐
相关产品推荐

