You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java中GZIP压缩字节写入文件后无法正常解压缩问题求助

问题:提取GZIP压缩负载并解压缩时抛出格式异常

我有格式为 A|B|A_VERY_LONG_STRING_THAT_WILL_BE_COMPRESSED|C|D 的字符串,按竖线|分割存入数组result[]:

result[0]=A;
result[1]=B;
result[2]=A_VERY_LONG_STRING_THAT_WILL_BE_COMPRESSED;
result[3]=C;
result[4]=D

使用以下方法压缩result[2]:

public static byte[] compressUsingStream(String payload) {

    try (ByteArrayOutputStream byteArrayOutputStream = new ByteArrayOutputStream();
         GZIPOutputStream gzipOutputStream = new GZIPOutputStream(byteArrayOutputStream)) {

        gzipOutputStream.write(payload.getBytes("UTF-8"));

        gzipOutputStream.finish();
        gzipOutputStream.close();

        return byteArrayOutputStream.toByteArray();

    } catch (IOException e) {
        throw new RuntimeException(e);
    }
}

调用方式:

byte[] compressedPayloadAsBytes = PayloadCompressionDecompression.compressUsingStream(result2);

将其他元素转字节数组后与压缩字节拼接:

byte[] finalArray = concatAll(result[0].getBytes(), 
"|".getBytes(), 
result[1].getBytes(), 
"|".getBytes(), 
compressedPayloadAsBytes, 
"|".getBytes(), 
result[3].getBytes(), 
"|".getBytes(), 
result[4].getBytes());

写入文件:

Path path = Path.of(file);
Files.write(path, finalArray);

读取文件时执行:

byte[]  allBytesFromFile = Files.readAllBytes(path);
String recordWithCompressedPayload = new String(allBytesFromFile);

截取压缩负载部分:

int payloadStart = StringUtils.ordinalIndexOf(recordWithCompressedPayload, "|", 2);
int payloadEnd = StringUtils.lastOrdinalIndexOf(recordWithCompressedPayload, "|", 2);

String compressedPayloadAsStr = recordWithCompressedPayload.substring(payloadStart+1, payloadEnd);

但将compressedPayloadAsStr传入解压缩方法时抛出异常java.lang.RuntimeException: java.util.zip.ZipException: Not in GZIP format:

public static String deCompressUsingStream(byte[] compressedPayload) {

    try (GZIPInputStream gzipInputStream = new GZIPInputStream(new ByteArrayInputStream(compressedPayload))) {

        final StringWriter stringWriter = new StringWriter();
        IOUtils.copy(gzipInputStream, stringWriter, UTF_8);
        gzipInputStream.close();
        return stringWriter.toString();
    } catch (IOException e) {
        throw new RuntimeException(e);
    }
}

调用方式:PayloadCompressionDecompression.deCompressUsingStream(compressedPayloadAsStr.getBytes())

需要解决如何正确从文件提取压缩负载并传入deCompressUsingStream()方法的问题。


问题根源

核心问题在于GZIP压缩后的字节数组是二进制数据,不能直接转成字符串再转回字节数组。二进制数据包含大量非UTF-8可打印字符,转成字符串时会出现编码丢失、字符替换的情况,导致转回的字节数组和原压缩字节不一致,GZIP无法识别格式。

同时,用字符串分割的方式提取压缩负载本身就不可靠——压缩后的二进制数据中可能包含|字符,会导致分割位置判断错误。

解决方案

需要直接操作字节数组定位压缩负载的位置,避免二进制和字符串的相互转换:

方案1:写入文件时记录各部分长度(推荐)

在写入文件时,先写入各字段的长度信息,再写入字段内容,这样读取时可以按长度精准提取:

// 写入时:先写各部分的长度(用固定4字节int存储),再写内容
ByteArrayOutputStream baos = new ByteArrayOutputStream();
DataOutputStream dos = new DataOutputStream(baos);
StandardCharsets utf8 = StandardCharsets.UTF_8;

// 写入result[0]的长度和内容
byte[] part0 = result[0].getBytes(utf8);
dos.writeInt(part0.length);
dos.write(part0);

// 写入result[1]的长度和内容
byte[] part1 = result[1].getBytes(utf8);
dos.writeInt(part1.length);
dos.write(part1);

// 写入压缩后的负载长度和内容
dos.writeInt(compressedPayloadAsBytes.length);
dos.write(compressedPayloadAsBytes);

// 写入result[3]的长度和内容
byte[] part3 = result[3].getBytes(utf8);
dos.writeInt(part3.length);
dos.write(part3);

// 写入result[4]的长度和内容
byte[] part4 = result[4].getBytes(utf8);
dos.writeInt(part4.length);
dos.write(part4);

// 写入文件
Files.write(path, baos.toByteArray());

读取时按长度提取:

DataInputStream dis = new DataInputStream(new ByteArrayInputStream(allBytesFromFile));
StandardCharsets utf8 = StandardCharsets.UTF_8;

// 读取result[0]
int len0 = dis.readInt();
byte[] part0 = new byte[len0];
dis.readFully(part0);
String result0 = new String(part0, utf8);

// 读取result[1]
int len1 = dis.readInt();
byte[] part1 = new byte[len1];
dis.readFully(part1);
String result1 = new String(part1, utf8);

// 读取压缩负载
int payloadLen = dis.readInt();
byte[] compressedPayload = new byte[payloadLen];
dis.readFully(compressedPayload);

// 读取result[3]和result[4]
int len3 = dis.readInt();
byte[] part3 = new byte[len3];
dis.readFully(part3);
String result3 = new String(part3, utf8);

int len4 = dis.readInt();
byte[] part4 = new byte[len4];
dis.readFully(part4);
String result4 = new String(part4, utf8);

// 解压缩
String originalPayload = PayloadCompressionDecompression.deCompressUsingStream(compressedPayload);

方案2:直接在字节数组中定位分隔符(不推荐)

如果不能修改写入格式,需要在字节数组中定位|的位置,而不是转成字符串:

byte[] delimiter = "|".getBytes(StandardCharsets.UTF_8);
int delimiterLen = delimiter.length;

// 找第2个|的位置(跳过前两个字段)
int count = 0;
int payloadStart = 0;
for (int i = 0; i <= allBytesFromFile.length - delimiterLen; i++) {
    boolean match = true;
    for (int j = 0; j < delimiterLen; j++) {
        if (allBytesFromFile[i + j] != delimiter[j]) {
            match = false;
            break;
        }
    }
    if (match) {
        count++;
        if (count == 2) {
            payloadStart = i + delimiterLen;
            break;
        }
    }
}

// 找倒数第2个|的位置
count = 0;
int payloadEnd = allBytesFromFile.length;
for (int i = allBytesFromFile.length - delimiterLen; i >= 0; i--) {
    boolean match = true;
    for (int j = 0; j < delimiterLen; j++) {
        if (allBytesFromFile[i + j] != delimiter[j]) {
            match = false;
            break;
        }
    }
    if (match) {
        count++;
        if (count == 2) {
            payloadEnd = i;
            break;
        }
    }
}

// 提取压缩负载字节数组
byte[] compressedPayload = Arrays.copyOfRange(allBytesFromFile, payloadStart, payloadEnd);

// 解压缩
String originalPayload = PayloadCompressionDecompression.deCompressUsingStream(compressedPayload);

注意:这种方法依然存在风险——如果压缩后的二进制数据中包含|的字节序列,会导致定位错误,因此优先推荐第一种记录长度的方案。


内容的提问来源于stack exchange,提问作者Raj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 21:15:43