使用Avro Reflect工具序列化Scala样例类报错,求解决方案
Scala样例类Avro序列化反序列化问题解决
问题描述
尝试使用Avro的ReflectDatumReader/ReflectDatumWriter对Scala样例类进行序列化反序列化,代码添加了默认参数但仍报错,错误提示找不到MyRecord.<init>()无参构造方法,核心堆栈信息如下:
Exception in thread "main" java.lang.RuntimeException: java.lang.RuntimeException: java.lang.NoSuchMethodException: MyRecord.<init>() at org.apache.avro.specific.SpecificData.newInstance(SpecificData.java:473) ... Caused by: java.lang.NoSuchMethodException: MyRecord.<init>() at java.base/java.lang.Class.getConstructor0(Class.java:3349) ...
原代码:
import org.apache.avro.io.{DecoderFactory, EncoderFactory} import org.apache.avro.reflect.{ReflectDatumReader, ReflectDatumWriter} import java.io.ByteArrayOutputStream import java.util.UUID case class MyRecord( string: String ="", bool: Boolean = false, bigInt: BigInt = 0, bigDecimal: BigDecimal = 0, ) object AvroEncodingDemoApp extends App { val parser = new org.apache.avro.Schema.Parser() val a = new MyRecord(string = "???", bool = false, bigInt = BigInt.long2bigInt(1), bigDecimal = BigDecimal.decimal(5)) val avroSchema = parser.parse( """ |{ | "type": "record", | "name": "MyRecord", | "fields": [{ | "name": "string", | "type": "string" | }, { | "name": "bool", | "type": "boolean" | }, { | "name": "bigInt", | "type": { | "type": "long", | "precision": 24, | "scale": 24 | } | }, { | "name": "bigDecimal", | "type": { | "type": "double", | "logicalType": "decimal", | "precision": 48, | "scale": 24 | } | }] |} |""".stripMargin) val writer = new ReflectDatumWriter[MyRecord](avroSchema) val boaStream = new ByteArrayOutputStream() val jsonEncoder = EncoderFactory.get.jsonEncoder(avroSchema, boaStream) writer.write(a, jsonEncoder) jsonEncoder.flush() val reader = new ReflectDatumReader[MyRecord](avroSchema) val jsonDecoder = DecoderFactory.get().jsonDecoder(avroSchema, new String(boaStream.toByteArray)) val output = reader.read(null, jsonDecoder) println(output) }
错误原因
- 无参构造器缺失:Scala样例类即使给所有字段加了默认参数,编译后也不会自动生成Java风格的无参构造方法,而Avro默认的
ReflectData会尝试调用无参构造器创建实例。 - Schema类型定义错误:
BigInt字段不需要额外的precision和scale配置,直接用long类型即可(若值超出long范围则用bytes)。BigDecimal用double类型会丢失精度,正确做法是用bytes类型配合decimal逻辑类型。
解决方案
方案一:使用Scala专用ReflectData(推荐)
Avro提供了ScalaReflectData,专门处理Scala特性(比如样例类的apply方法、默认参数),无需手动添加无参构造器。
依赖配置(以sbt为例)
确保添加avro-scala依赖,版本与avro-core一致:
libraryDependencies ++= Seq( "org.apache.avro" % "avro-core" % "1.11.3", "org.apache.avro" % "avro-scala" % "1.11.3" )
完整示例代码
import org.apache.avro.io.{DecoderFactory, EncoderFactory} import org.apache.avro.reflect.{ReflectDatumReader, ReflectDatumWriter, ScalaReflectData} import java.io.ByteArrayOutputStream // 普通样例类,无需额外构造器 case class MyRecord( string: String = "", bool: Boolean = false, bigInt: BigInt = 0, bigDecimal: BigDecimal = 0 ) object AvroEncodingDemoApp extends App { // 使用ScalaReflectData处理Scala类型 val scalaReflectData = ScalaReflectData.get() // 从样例类自动生成Schema,避免手动编写出错 val avroSchema = scalaReflectData.getSchema(classOf[MyRecord]) println("自动生成的Schema:") println(avroSchema.toString(true)) // 待序列化的实例 val sourceRecord = MyRecord( string = "测试内容", bool = true, bigInt = BigInt(987654), bigDecimal = BigDecimal("1234.5678") ) // 序列化流程 val writer = new ReflectDatumWriter[MyRecord](avroSchema, scalaReflectData) val outputStream = new ByteArrayOutputStream() val jsonEncoder = EncoderFactory.get.jsonEncoder(avroSchema, outputStream) writer.write(sourceRecord, jsonEncoder) jsonEncoder.flush() val serializedJson = outputStream.toString("UTF-8") println("\n序列化后的JSON:") println(serializedJson) // 反序列化流程 val reader = new ReflectDatumReader[MyRecord](avroSchema, scalaReflectData) val jsonDecoder = DecoderFactory.get().jsonDecoder(avroSchema, serializedJson) val deserializedRecord = reader.read(null, jsonDecoder) println("\n反序列化后的实例:") println(deserializedRecord) }
方案二:手动添加无参构造器
如果不想依赖avro-scala,可以给样例类手动添加无参构造方法,同时修正Schema类型:
修正后的样例类
case class MyRecord( string: String = "", bool: Boolean = false, bigInt: BigInt = 0, bigDecimal: BigDecimal = 0 ) { // 手动添加无参构造方法,对应Java风格的无参构造 def this() = this("", false, 0, 0) }
修正后的Schema
{ "type": "record", "name": "MyRecord", "namespace": "com.yourpackage", // 需与样例类包名一致 "fields": [ { "name": "string", "type": "string", "default": "" }, { "name": "bool", "type": "boolean", "default": false }, { "name": "bigInt", "type": "long", "default": 0 }, { "name": "bigDecimal", "type": { "type": "bytes", "logicalType": "decimal", "precision": 18, "scale": 6 }, "default": "\u0000" } ] }
注意事项
- 若使用手动编写的Schema,必须保证
namespace与样例类的包名一致,否则Avro反射时无法找到对应类。 BigDecimal使用bytes类型配合decimal逻辑类型能保证精度,避免double带来的精度丢失问题。
内容的提问来源于stack exchange,提问作者user51
相关产品推荐
相关产品推荐

