Avro与Protobuf序列化性能对比及Avro代码正确性问询
Hey there, let's dig into why your Avro serialization throughput is lagging so far behind Protobuf's. Looking at your code, there are several critical inefficiencies that are almost certainly dragging down performance—let's break them down and fix them:
Key Issues in Your Current Code
1. Frequent Allocation of ByteArrayOutputStream
Every call to serialize() creates a new ByteArrayOutputStream instance. For high-throughput benchmarks, this leads to massive garbage collection (GC) overhead, which is a major performance killer. Protobuf handles buffer reuse internally, so this is a big gap between your two implementations.
2. Thread-Unsafe Encoder Reuse & Wasted Buffering
Your encoder is an instance variable, which is not thread-safe—JMH runs benchmarks in multiple threads by default, so concurrent access to this shared encoder will cause race conditions, corrupted data, and slowdowns. Additionally, since you're passing a new ByteArrayOutputStream each time, the encoder's internal buffer isn't being reused effectively.
3. Unnecessary Object Creation in createAvro()
You're creating a new AvroGeneratedClass.Builder and AvroGeneratedClass instance for every serialization call. This adds extra GC pressure and constructor overhead that Protobuf avoids by allowing object reuse.
Fixed Implementation
Here's a revised version of your code that addresses these issues, focusing on reuse and thread safety:
public final class AvroSerialization { // Use ThreadLocal to ensure thread safety and reuse across calls private final ThreadLocal<ByteArrayOutputStream> outThreadLocal = ThreadLocal.withInitial(() -> new ByteArrayOutputStream(256)); // Match your ~200 byte payload private final ThreadLocal<BinaryEncoder> encoderThreadLocal = ThreadLocal.withInitial(() -> EncoderFactory.get().binaryEncoder(outThreadLocal.get(), null)); private final SpecificDatumWriter<AvroGeneratedClass> writer; private final ThreadLocal<AvroGeneratedClass> avroObjectThreadLocal = ThreadLocal.withInitial(AvroGeneratedClass::new); public AvroSerialization() { this.writer = new SpecificDatumWriter<>(AvroGeneratedClass.class); } public final byte[] serialize(MyDataObject data) { ByteArrayOutputStream out = outThreadLocal.get(); out.reset(); // Reuse the stream instead of creating a new one BinaryEncoder encoder = encoderThreadLocal.get(); encoder = EncoderFactory.get().binaryEncoder(out, encoder); // Reuse encoder with reset stream AvroGeneratedClass avroData = avroObjectThreadLocal.get(); populateAvro(avroData, data); // Reuse existing object instead of creating new try { writer.write(avroData, encoder); encoder.flush(); return out.toByteArray(); } catch (IOException e) { throw new RuntimeException("Avro serialization failed", e); } } // Reuse existing Avro object instead of creating new instances private void populateAvro(AvroGeneratedClass avroData, MyDataObject data) { avroData.clear(); // Reset object to default state avroData.setXXX(data.getXXX()); avroData.setXXX(data.getXXX()); // ... set other fields } }
Explanation of Changes
- ThreadLocal Reuse: All stateful objects (
ByteArrayOutputStream,BinaryEncoder,AvroGeneratedClass) are stored inThreadLocalvariables. This ensures thread safety and allows reuse across multiple serialization calls in the same thread, eliminating most GC overhead. - Stream Reset: Instead of creating a new
ByteArrayOutputStream, we callreset()to clear the existing stream's content—this is far cheaper than allocating a new object. - Object Reuse: The
AvroGeneratedClassinstance is reused across calls, withclear()resetting it to a clean state before populating new data. This avoids the cost of creating and garbage collecting thousands of objects per second. - Correct Encoder Reuse: By reusing the encoder with the reset stream, we leverage its internal buffer, reducing memory allocations.
Additional Performance Tips
- Use
DirectBinaryEncoder: For even better performance, tryEncoderFactory.get().directBinaryEncoder(out, encoder)—it uses direct buffers which can reduce copy overhead. - Tune JMH Parameters: Ensure your JMH benchmark uses sufficient warm-up iterations (e.g.,
-wi 5 -i 10) and runs in throughput mode (-bm thrpt), matching your Protobuf test setup. - Check Schema Efficiency: Make sure your Avro schema doesn't include unnecessary fields or complex types that add serialization overhead.
After making these changes, you should see a significant improvement in Avro's throughput—while it might still not match Protobuf's raw speed, it should be much closer than your current 80k ops/sec.
内容的提问来源于stack exchange,提问作者CaptainHastings

