Avro格式Kafka消息创建报错:Expected start-union. Got VALUE_STRING 求助
问题描述
我正在使用以下Avro Schema测试Avro格式数据:
{ "namespace": "de.leyland.audit", "type": "record", "name": "AuditDataChangeleyland", "fields": [ {"name": "machineID", "type": "string"}, {"name": "correlationId", "type": "string"}, {"name": "timestamp", "type": "long", "logicalType": "timestamp-millis"}, {"name": "machinescreen","type":{"type": "enum", "name": "BlockScreen", "symbols": ["NO","YES"]}}, {"name": "entitlements", "type": ["null",{ "type": "array", "items": { "name": "entitlements", "type": "record", "fields": [ { "name": "name", "type": "string" } ] } }]} ] }
尝试基于该Schema创建供Kafka消费者使用的JSON格式测试数据时出现报错,测试用JSON如下:
{ "machineID": "ahdzeha46", "correlationId": "473363621", "timestamp": "2021-08-09T12:20:15", "machinescreen": "NO", "correlationId": "corr-473363621", "entitlements": [ { "entitlements": "machinerepair" } ] }
报错信息:
Caused by: org.apache.avro.AvroTypeException: Expected start-union. Got VALUE_STRING
请问如何修正并生成正确的测试数据?
错误分析与修正方案
你的测试数据存在4个关键问题,逐一修正即可解决报错:
- Union类型处理错误:
entitlements是["null", array]的Union类型,Avro JSON要求Union类型数据必须用带类型标识的对象包裹(null除外),你直接传数组不符合规则。 - 数组内部结构不匹配:Schema中数组的items是含
name字段的record,但你用了entitlements作为键名,结构完全不符。 - 时间戳类型错误:
timestamp定义为毫秒级时间戳(long类型),不能传ISO格式字符串,需传入数值。 - 重复字段冗余:测试数据重复定义了
correlationId,需删除其中一个。
修正后的正确测试数据
如果要传入有值的entitlements,正确格式如下:
{ "machineID": "ahdzeha46", "correlationId": "corr-473363621", "timestamp": 1628506815000, "machinescreen": "NO", "entitlements": { "array": [ { "name": "machinerepair" } ] } }
如果要传入null值的entitlements,可以写成:
{ "machineID": "ahdzeha46", "correlationId": "corr-473363621", "timestamp": 1628506815000, "machinescreen": "NO", "entitlements": null }
内容的提问来源于stack exchange,提问作者data2quest
相关产品推荐
相关产品推荐

