EHCache中Kryo序列化器占用空间大于默认序列化器的原因排查
我尝试用Kryo序列化器将Employee对象存入EHCache堆外层级,却发现空间占用比默认序列化器大很多,具体统计数据如下:
使用Kryo时的堆外层级统计:
tierStats=[TierStats(tierName=OffHeap, allocatedByteSize=6094848, occupiedByteSize=4184000, evictions=0, expirations=0, hits=0, misses=0, mappings=1000, puts=0, removals=0)])
使用默认序列化器时的堆外层级统计:
tierStats=[TierStats(tierName=OffHeap, allocatedByteSize=4259840, occupiedByteSize=119200, evictions=0, expirations=0, hits=0, misses=0, mappings=1000, puts=0, removals=0)])
另外,我用Kryo单独序列化同一组1000个对象到文件,文件大小仅约4KB,远小于EHCache中统计的occupiedByteSize。
使用的版本:
implementation 'org.ehcache:ehcache:3.10.0' implementation 'com.esotericsoftware:kryo:5.6.0'
代码片段
Employee类
public class Employee implements Serializable { String name; int id; public Employee(String name, int id) { this.name = name; this.id = id; } }
Kryo序列化器实现
public class EmployeeKryoSerializer implements Serializer<Employee> { private static final Kryo kryo = new Kryo(); public EmployeeKryoSerializer(ClassLoader loader) { kryo.register(Employee.class); } @Override public ByteBuffer serialize(Employee object) throws SerializerException { Output output = new Output(new ByteArrayOutputStream()); kryo.writeObject(output, object); output.close(); return ByteBuffer.wrap(output.getBuffer()); } @Override public Employee read(ByteBuffer binary) throws SerializerException { Input input = new Input(new ByteBufferInputStream(binary)); return kryo.readObject(input, Employee.class); } @Override public boolean equals(Employee object, ByteBuffer binary) throws ClassNotFoundException, SerializerException { return object.equals(read(binary)); } }
测试代码
public class KryoSerializationTest { public static void main(String[] args) throws IOException { Employee test = new Employee("TestName", 1); testNormalKryoDummyObject(test); testEHCacheDummyObject(test); } private static void testNormalKryoDummyObject( Employee test) throws IOException { Kryo kryo = new Kryo(); kryo.register(Employee.class); Output output = new Output(new ByteArrayOutputStream()); for (int i = 0; i < 1000; i++) { kryo.writeObject(output, test); } output.close(); FileOutputStream fileOutputStream = new FileOutputStream("kryo_employee.bin"); fileOutputStream.write(output.getBuffer()); fileOutputStream.flush(); fileOutputStream.close(); } private static void testEHCacheDummyObject(Employee test){ StatisticsService odStatisticsService = new DefaultStatisticsService(); CacheManager cacheManager = CacheManagerBuilder.newCacheManagerBuilder() .using(odStatisticsService) .build(true); Cache<String, Employee> odPairCache = cacheManager.createCache("TEST_EH_CACHE", CacheConfigurationBuilder.newCacheConfigurationBuilder( String.class, Employee.class, ResourcePoolsBuilder.newResourcePoolsBuilder().offheap(100, MemoryUnit.MB).build()) .withValueSerializer(EmployeeKryoSerializer.class) .build()); for (int i = 0; i < 1000; i++) { odPairCache.put("XYZ"+i+"-"+"ABC",test); printEHStatistic("TEST_EH_CACHE", odStatisticsService); } } private static void printEHStatistic(String cacheName, StatisticsService odStatisticsService) { CacheStatistics cacheStatistics = odStatisticsService.getCacheStatistics(cacheName); EHCacheStatistic ehCacheStatistic = EHCacheStatistic.builder() .cacheName(cacheName) .cacheMissPercentage(cacheStatistics.getCacheMissPercentage()) .cacheEvictions(cacheStatistics.getCacheEvictions()) .cacheExpirations(cacheStatistics.getCacheExpirations()) .cacheGets(cacheStatistics.getCacheGets()) .cacheMisses(cacheStatistics.getCacheMisses()) .cacheRemovals(cacheStatistics.getCacheRemovals()) .cachePuts(cacheStatistics.getCachePuts()) .cacheHits(cacheStatistics.getCacheHits()) .cacheHitPercentage(cacheStatistics.getCacheHitPercentage()) .tierStats(new ArrayList<>()) .build(); cacheStatistics .getTierStatistics() .forEach( (tierName, tierStats) -> ehCacheStatistic .getTierStats() .add( TierStats.builder() .tierName(tierName) .allocatedByteSize(tierStats.getAllocatedByteSize()) .occupiedByteSize(tierStats.getOccupiedByteSize()) .evictions(tierStats.getEvictions()) .expirations(tierStats.getExpirations()) .hits(tierStats.getHits()) .misses(tierStats.getMisses()) .mappings(tierStats.getMappings()) .puts(tierStats.getPuts()) .removals(tierStats.getRemovals()) .build())); System.out.println(ehCacheStatistic.toString()); } }
原因分析
Kryo序列化器的ByteBuffer处理不当
当前实现中,output.getBuffer()返回的是Kryo Output的整个初始缓冲区(默认大小2048字节),而非实际序列化后的有效数据长度。EHCache会将这个ByteBuffer的全部大小计入堆外占用空间,哪怕其中大部分是未使用的空白字节。独立测试中是把1000个对象写入同一个输出流,Kryo会自动扩容并复用缓冲区,最终有效数据仅4KB;但EHCache里每个对象单独序列化,每个都带着2048字节的缓冲区,1000个就是2MB左右,再加上EHCache的条目元数据,就会达到统计的4MB。EHCache默认序列化器的优化机制
EHCache默认序列化器基于Java序列化做了针对性优化:- 对重复对象启用引用共享,1000个相同的Employee对象只会实际存储一次,其他条目存储引用,因此占用空间极小(约116KB)。
- 堆外存储时会做内存块对齐、复用的优化,减少额外开销。
自定义Kryo序列化器绕过了这些优化,每个对象都被单独序列化存储,没有引用共享。
equals方法的低效实现
当前equals方法会直接反序列化整个对象再比较,不仅影响性能,还可能导致EHCache在内部做一致性检查时产生额外临时对象,间接增加空间占用。
解决方案
修正Kryo序列化的ByteBuffer生成逻辑
替换output.getBuffer()为output.toBytes(),该方法返回实际序列化后的有效字节数组,而非整个缓冲区;同时可以指定更小的初始缓冲区大小(匹配对象实际大小):@Override public ByteBuffer serialize(Employee object) throws SerializerException { Output output = new Output(100); // 初始化更小的缓冲区 kryo.writeObject(output, object); output.close(); return ByteBuffer.wrap(output.toBytes()); }启用Kryo的引用共享
给Kryo实例开启引用共享,重复对象只会序列化一次:private static final Kryo kryo = new Kryo(); static { kryo.setReferences(true); // 开启引用共享 kryo.register(Employee.class); }优化equals方法
避免通过反序列化比较对象,直接提取ByteBuffer中的关键字段进行比较:@Override public boolean equals(Employee object, ByteBuffer binary) throws ClassNotFoundException, SerializerException { Input input = new Input(new ByteBufferInputStream(binary.duplicate())); Employee cached = kryo.readObject(input, Employee.class); return Objects.equals(object.id, cached.id) && Objects.equals(object.name, cached.name); }
内容的提问来源于stack exchange,提问作者nil186007

