You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用zlib将压缩缓冲区写入gzip兼容文件时解压报错,求排查

问题描述

我希望将缓冲区中的大量数据压缩后写入gzip兼容文件,这么做是因为有多线程可并行压缩各自的数据,仅在写入公共输出文件时需要加锁。我基于zlib.h文档编写了测试代码,但尝试解压输出文件时出现错误:gzip: test.gz: unexpected end of file。请问可能是什么问题导致的?

测试代码如下:

#include <cassert>
#include <fstream>
#include <string.h>
#include <zlib.h>


int main()
{

    char compress_in[50] = "I waaaaaaaant tooooooo beeeee compressssssed";
    char compress_out[100];

    z_stream bufstream;
    bufstream.zalloc = Z_NULL;
    bufstream.zfree = Z_NULL;
    bufstream.opaque = Z_NULL;

    bufstream.avail_in = ( uInt )strlen(compress_in) + 1;
    bufstream.next_in = ( Bytef * ) compress_in;
    bufstream.avail_out = ( uInt )sizeof( compress_out );
    bufstream.next_out = ( Bytef * ) compress_out;

    int res = deflateInit2( &bufstream, Z_DEFAULT_COMPRESSION, Z_DEFLATED, 15 + 16, 8, Z_DEFAULT_STRATEGY );
    assert ( res == Z_OK );

    res = deflate( &bufstream, Z_FINISH );
    assert( res == Z_STREAM_END );

    deflateEnd( &bufstream );

    std::ofstream outfile( "test.gz", std::ios::binary | std::ios::out );
    outfile.write( compress_out, strlen( compress_out ) + 1 );
    outfile.close();

    return 0;
}
问题原因

核心问题出在压缩后的数据长度计算错误:

  • 你用strlen(compress_out) + 1来获取压缩后的数据长度,但压缩后的数据是二进制流,里面大概率包含\0字符,strlen会在第一个\0处停止计数,导致实际写入文件的字节数远小于真实的压缩数据长度,缺失的部分让gzip判定文件不完整,从而抛出"unexpected end of file"错误。
  • zlib的z_stream结构体自带的total_out字段会准确记录压缩后输出的总字节数,这才是写入文件时应该使用的长度。

另外还有个潜在风险:当前测试用的compress_out缓冲区大小是100,虽然对测试数据足够,但处理大量数据时可能出现缓冲区不足的情况,不过这不是当前报错的直接原因。

修复后的代码
#include <cassert>
#include <fstream>
#include <string.h>
#include <zlib.h>


int main()
{

    char compress_in[50] = "I waaaaaaaant tooooooo beeeee compressssssed";
    char compress_out[100];

    z_stream bufstream;
    bufstream.zalloc = Z_NULL;
    bufstream.zfree = Z_NULL;
    bufstream.opaque = Z_NULL;

    bufstream.avail_in = ( uInt )strlen(compress_in) + 1;
    bufstream.next_in = ( Bytef * ) compress_in;
    bufstream.avail_out = ( uInt )sizeof( compress_out );
    bufstream.next_out = ( Bytef * ) compress_out;

    int res = deflateInit2( &bufstream, Z_DEFAULT_COMPRESSION, Z_DEFLATED, 15 + 16, 8, Z_DEFAULT_STRATEGY );
    assert ( res == Z_OK );

    res = deflate( &bufstream, Z_FINISH );
    assert( res == Z_STREAM_END );

    // 获取真实的压缩数据总长度
    uLong compressed_len = bufstream.total_out;

    deflateEnd( &bufstream );

    std::ofstream outfile( "test.gz", std::ios::binary | std::ios::out );
    // 使用total_out记录的长度写入完整压缩数据
    outfile.write( compress_out, compressed_len );
    outfile.close();

    return 0;
}
多线程场景额外提示

针对你提到的多线程并行压缩后写入同一gzip文件的场景,注意这两点:

  • gzip格式支持多个独立的压缩块(称为"成员"),每个线程压缩的数据可以作为一个独立的gzip成员写入文件,不需要提前合并,只要写入时加锁保证顺序即可。
  • 每个线程必须独立初始化z_stream结构体,完整执行deflateInit2→deflate→deflateEnd流程,生成包含头、压缩数据、校验和的完整gzip成员,再写入文件,解压时gzip会自动处理多个成员。

内容的提问来源于stack exchange,提问作者Raghavendra P

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 11:45:35