使用Apache Avro写入Parquet时BigDecimal转ByteBuffer报错求助
解决Avro生成POJO写入Parquet时BigDecimal转ByteBuffer异常
问题定位
你遇到的java.math.BigDecimal cannot be cast to java.nio.ByteBuffer异常,本质是Parquet-Avro在处理Decimal逻辑类型时,未正确完成BigDecimal到ByteBuffer的自动转换,核心原因有两个:
- 依赖版本不兼容:
hadoop-core 1.2.1版本过于陈旧,与parquet-avro 1.11.1的Decimal逻辑类型支持不匹配 - ParquetWriter未配置Decimal逻辑类型的处理规则
解决方案
1. 替换陈旧的Hadoop依赖
将过时的hadoop-core替换为与当前Parquet-Avro版本兼容的Hadoop客户端依赖,示例如下:
<dependency> <groupId>org.apache.hadoop</groupId> <artifactId>hadoop-client</artifactId> <version>2.7.7</version> </dependency>
2. 配置ParquetWriter启用Decimal逻辑类型处理
在构建AvroParquetWriter时,添加配置项明确启用Decimal逻辑类型转换,确保Parquet能正确识别Avro定义的Decimal字段:
Configuration conf = new Configuration(); // 开启Avro兼容性模式,支持逻辑类型 conf.setBoolean(AvroReadSupport.AVRO_COMPATIBILITY, true); // 传入Avro写入Schema,明确字段类型规则 conf.set(ParquetAvroConstants.AVRO_WRITE_SCHEMA, employee.getSchema().toString());
3. 验证生成的POJO字段类型
确认Avro Maven插件生成的Employee类中,salary字段确实为BigDecimal类型(你已开启<enableDecimalLogicalType>true</enableDecimalLogicalType>,此步骤通常无需额外调整)。
完整修正后的核心代码
Main类代码
public static void main(String[] args) throws IOException { File outputParquet = new File("./output.parquet"); Files.deleteIfExists(outputParquet.toPath()); Employee employee = new Employee("john", "john.doe@mail.com", BigDecimal.TEN); Configuration conf = new Configuration(); conf.setBoolean(AvroReadSupport.AVRO_COMPATIBILITY, true); conf.set(ParquetAvroConstants.AVRO_WRITE_SCHEMA, employee.getSchema().toString()); ParquetWriter<Employee> writer = AvroParquetWriter.<Employee>builder(new Path(outputParquet.getAbsolutePath())) .withCompressionCodec(CompressionCodecName.SNAPPY) .withSchema(employee.getSchema()) .withConf(conf) .build(); try { writer.write(employee); } catch (Exception e) { e.printStackTrace(); } finally { writer.close(); } }
pom.xml核心依赖
<dependencies> <dependency> <groupId>org.apache.parquet</groupId> <artifactId>parquet-avro</artifactId> <version>1.11.1</version> </dependency> <dependency> <groupId>org.apache.avro</groupId> <artifactId>avro</artifactId> <version>1.11.1</version> </dependency> <dependency> <groupId>org.apache.hadoop</groupId> <artifactId>hadoop-client</artifactId> <version>2.7.7</version> </dependency> </dependencies>
内容的提问来源于stack exchange,提问作者HakanGuneser
相关产品推荐
相关产品推荐

