Java中使用GZIP BEST_COMPRESSION压缩后解压长度异常排查
GZIP压缩使用BEST_COMPRESSION级别后解压数据长度异常问题
我希望在将字符串存入数据库前对其进行压缩,已经掌握不使用BEST_COMPRESSION级别时的GZIP压缩方法,但为了获得更好的压缩效果,想要使用Deflater.BEST_COMPRESSION级别。测试时,对字符串“hello”(原长度5)进行压缩后解压,结果显示解压后长度为4,不清楚问题出在哪里。
相关代码如下:
class ZIP extends GZIPOutputStream { public ZIP(OutputStream out) throws IOException { super(out); def.setLevel(Deflater.BEST_COMPRESSION); } } public class Main { public static byte[] in(byte[] b) throws Exception { byte[] readBuffer = new byte[5000]; ByteArrayInputStream in = new ByteArrayInputStream(b); GZIPInputStream zipIn = new GZIPInputStream(in); b = Arrays.copyOf(readBuffer, zipIn.read(readBuffer, 0, readBuffer.length)); in.close(); zipIn.close(); return b; } public static byte[] out(byte[] b) throws Exception { ByteArrayOutputStream arrayOutputStream = new ByteArrayOutputStream(); ZIP out = new ZIP(arrayOutputStream); new ObjectOutputStream(out).write(b); out.finish(); arrayOutputStream.flush(); return arrayOutputStream.toByteArray(); } public static void main(String[] a) throws Exception { String message = "hello"; System.out.println("orig: " + message.length()); byte[] out = out(message.getBytes(StandardCharsets.UTF_8)); byte[] in = in(out); System.out.println("decompressed: " + new String(in, StandardCharsets.UTF_8).length()); } }
测试输出结果:
orig: 5 decompressed: 4
问题原因
- 解压时未完整读取数据:你只调用了一次
zipIn.read(),但read()方法无法保证一次性读取完所有解压后的内容。当使用BEST_COMPRESSION级别时,压缩后的数据流结构可能导致第一次读取仅返回部分字节,剩余数据未被读取,最终造成解压后数组长度不足。 - 冗余的ObjectOutputStream:压缩时使用
ObjectOutputStream完全没必要,它会额外写入序列化魔数和版本号等头部信息,既增加了压缩体积,也可能引入不必要的解析问题。
修正后的代码
import java.io.*; import java.nio.charset.StandardCharsets; import java.util.zip.Deflater; import java.util.zip.GZIPInputStream; import java.util.zip.GZIPOutputStream; class ZIP extends GZIPOutputStream { public ZIP(OutputStream out) throws IOException { super(out); def.setLevel(Deflater.BEST_COMPRESSION); } } public class Main { public static byte[] decompress(byte[] compressedData) throws IOException { ByteArrayInputStream in = new ByteArrayInputStream(compressedData); GZIPInputStream zipIn = new GZIPInputStream(in); ByteArrayOutputStream out = new ByteArrayOutputStream(); byte[] readBuffer = new byte[5000]; int bytesRead; // 循环读取直到流结束,确保获取所有解压数据 while ((bytesRead = zipIn.read(readBuffer)) != -1) { out.write(readBuffer, 0, bytesRead); } zipIn.close(); in.close(); return out.toByteArray(); } public static byte[] compress(byte[] rawData) throws IOException { ByteArrayOutputStream arrayOutputStream = new ByteArrayOutputStream(); ZIP out = new ZIP(arrayOutputStream); // 直接写入原始字节数组,移除冗余的ObjectOutputStream out.write(rawData); out.finish(); arrayOutputStream.flush(); return arrayOutputStream.toByteArray(); } public static void main(String[] a) throws Exception { String message = "hello"; System.out.println("orig: " + message.length()); byte[] compressed = compress(message.getBytes(StandardCharsets.UTF_8)); byte[] decompressed = decompress(compressed); String result = new String(decompressed, StandardCharsets.UTF_8); System.out.println("decompressed: " + result.length()); System.out.println("decompressed content: " + result); } }
修正说明
- 循环读取解压数据:通过
while循环持续读取GZIPInputStream,直到read()返回-1(表示流已读完),确保所有解压后的数据都被收集。 - 移除ObjectOutputStream:直接通过
GZIPOutputStream写入原始字节数组,避免不必要的序列化头部,同时简化压缩逻辑。
修正后的测试输出:
orig: 5 decompressed: 5 decompressed content: hello
内容的提问来源于stack exchange,提问作者speedcubing
相关产品推荐
相关产品推荐

