You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Avro ProtobufDatumReader解码标准Protobuf数据至Avro POJO失败求助

问题

我正在Java中尝试解码Protobuf编码的数据,目标是将其序列化为Avro POJO。

使用的测试Protobuf文件(Sample.proto)如下:

syntax = "proto3";

option java_package = "protodecoder.decoder";

message SampleMsg {
    string firstname = 1;
    string lastname = 2;
}

该数据的十六进制编码为:

0a044a6f686e1203446f65

对应的字节数组为:

[10, 4, 74, 111, 104, 110, 18, 3, 68, 111, 101]

解码后预期结果:

{ "firstname": "John", "lastname": "Doe" }

用Protobuf库可以正常解码为Protobuf对象:

Sample.SampleMsg message = Sample.SampleMsg.parseFrom(protobufData);

但之后需要手动转换为Avro POJO,想找更便捷的方式。

尝试用Avro Protobuf库转换时,代码如下:

byte[] protobufData = new byte[] { 10, 4, 74, 111, 104, 110, 18, 3, 68, 111, 101 };
ProtobufDatumReader<Sample.SampleMsg> datumReader = new ProtobufDatumReader<>(Sample.SampleMsg.class);
GenericDatumReader<GenericRecord> genericDatumReader = new GenericDatumReader<GenericRecord>(datumReader.getSchema());
GenericRecord record = genericDatumReader.read(null, DecoderFactory.get().binaryDecoder(new ByteArrayInputStream(protobufData), null));

解码时抛出EOFException:

Exception in thread "main" java.io.EOFException
    at org.apache.avro.io.BinaryDecoder$InputStreamByteSource.readRaw(BinaryDecoder.java:855)
    at org.apache.avro.io.BinaryDecoder.doReadBytes(BinaryDecoder.java:372)
    at org.apache.avro.io.BinaryDecoder.readString(BinaryDecoder.java:289)
    at org.apache.avro.io.BinaryDecoder.readString(BinaryDecoder.java:298)
    at org.apache.avro.io.ResolvingDecoder.readString(ResolvingDecoder.java:219)
    at org.apache.avro.generic.GenericDatumReader.readString(GenericDatumReader.java:457)
    at org.apache.avro.generic.GenericDatumReader.readWithoutConversion(GenericDatumReader.java:192)
    at org.apache.avro.generic.GenericDatumReader.read(GenericDatumReader.java:161)
    at org.apache.avro.generic.GenericDatumReader.readField(GenericDatumReader.java:260)
    at org.apache.avro.generic.GenericDatumReader.readRecord(GenericDatumReader.java:248)
    at org.apache.avro.generic.GenericDatumReader.readWithoutConversion(GenericDatumReader.java:180)
    at org.apache.avro.generic.GenericDatumReader.read(GenericDatumReader.java:161)
    at org.apache.avro.generic.GenericDatumReader.read(GenericDatumReader.java:154)

当用Avro自身编码Protobuf对象时,代码如下:

Sample.SampleMsg.Builder builder = Sample.SampleMsg.newBuilder();
builder.setFirstname("John");
builder.setLastname("Doe");
Sample.SampleMsg todo = builder.build();

ProtobufDatumWriter<Sample.SampleMsg> datumWriter = new ProtobufDatumWriter<>(Sample.SampleMsg.class);
ByteArrayOutputStream os = new ByteArrayOutputStream();
Encoder e = EncoderFactory.get().binaryEncoder(os, null);
datumWriter.write(todo, e);
e.flush();

生成的字节数组为:

[8, 74, 111, 104, 110, 6, 68, 111, 101]

该数组可被Avro解码,但用Google Protobuf库解析时抛出异常:

Exception in thread "main" com.google.protobuf.InvalidProtocolBufferException$InvalidWireTypeException: Protocol message tag had invalid wire type.
    at com.google.protobuf.InvalidProtocolBufferException.invalidWireType(InvalidProtocolBufferException.java:142)
    at com.google.protobuf.UnknownFieldSet$Builder.mergeFieldFrom(UnknownFieldSet.java:526)
    at com.google.protobuf.GeneratedMessageV3$Builder.parseUnknownField(GeneratedMessageV3.java:887)
    at protodecoder.decoder.Sample$SampleMsg$Builder.mergeFrom(Sample.java:494)
    at protodecoder.decoder.Sample$SampleMsg$1.parsePartialFrom(Sample.java:688)
    at protodecoder.decoder.Sample$SampleMsg$1.parsePartialFrom(Sample.java:680)
    at com.google.protobuf.AbstractParser.parsePartialFrom(AbstractParser.java:158)
    at com.google.protobuf.AbstractParser.parseFrom(AbstractParser.java:191)
    at com.google.protobuf.AbstractParser.parseFrom(AbstractParser.java:203)
    at com.google.protobuf.AbstractParser.parseFrom(AbstractParser.java:208)
    at com.google.protobuf.AbstractParser.parseFrom(AbstractParser.java:48)

请问:Avro ProtobufDatumReader是否仅支持解码由Avro库编码的Protobuf数据?还是我操作有误?若确实如此,是否有其他库可实现将标准Protobuf编码数据转换为Avro POJO的需求?

回答

关于Avro ProtobufDatumReader的兼容性

Avro的ProtobufDatumReader/ProtobufDatumWriter不是用来直接处理标准Protobuf二进制编码数据的。它们的作用是:基于Protobuf的定义生成对应的Avro Schema,然后将Protobuf对象序列化为Avro二进制格式,或者从Avro二进制格式反序列化为Protobuf对象。

你遇到的问题根源就在这里:

  • 用标准Protobuf编码的数据是Protobuf二进制格式,而GenericDatumReader期望读取的是Avro二进制格式,所以会抛出EOFException;
  • 用ProtobufDatumWriter生成的是Avro二进制格式,自然无法被标准Protobuf库解析,所以抛出WireType异常。

正确的转换路径

如果要把标准Protobuf编码的数据转成Avro POJO,正确的流程是:

  1. 先用标准Protobuf库把二进制数据解析成Protobuf对象;
  2. 再把Protobuf对象转换成Avro对象。

自动化转换的方案

如果你不想手动编写转换代码,可以考虑以下两种方式:

  1. 使用Apache Avro的Protobuf转换工具类:
    Avro提供了ProtobufData类,可以直接将Protobuf对象转换成Avro的GenericRecord,示例代码:
    byte[] protobufData = new byte[] { 10, 4, 74, 111, 104, 110, 18, 3, 68, 111, 101 };
    // 第一步:解析成Protobuf对象
    Sample.SampleMsg protoMsg = Sample.SampleMsg.parseFrom(protobufData);
    // 第二步:转换为Avro GenericRecord
    ProtobufData protobufDataHelper = new ProtobufData();
    GenericRecord avroRecord = protobufDataHelper.toRecord(protoMsg);
    // 如果需要转换成具体的Avro POJO,还可以用SpecificDatumReader进一步转换
    
  2. 使用代码生成工具:
    基于Protobuf定义生成对应的Avro POJO和转换代码,比如:
    • 用Avro的protoc插件(avro-protobuf),从.proto文件生成Avro Schema和对应的Java类;
    • 或者使用MapStruct这类映射框架,自动生成Protobuf对象到Avro POJO的映射代码,减少手动编写的工作量。

内容的提问来源于stack exchange,提问作者knightWizard

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 00:54:53