解析Git Pack时zlib报incorrect header check错误的求助
Git Pack文件解析报错:zlib返回Z_DATA_ERROR或invalid stored block lengths
问题背景
我通过POST请求获取Git Pack文件,请求地址为https://github.com/codecrafters-io/git-sample-3/git-upload-pack,请求数据如下:
0032want 23f0bc3b5c7c3108e41c448f01a3db31e7064bbb 00000009done
获取到的内容存储在std::string pack中,解析代码如下:
pack = pack.substr(20, pack.length() - 40); // skip 8 byte http header, 12 byte pack header, 20 byte pack trailer int type; int pos = 0; std::string lengthstr; int length = 0; type = (pack[pos] & 112) >> 4; // 112 is 11100000 so this gets the first 3 bits length = length | (pack[pos] & 0x0F); // take the last 4 bits if (pack[pos] & 0x80) { // if the type bit starts with 0 then the last 4 bits are simply the length pos++; while (pack[pos] & 0x80) { // while leftmost bit is 1 length = length << 7; length = length | (pack[pos] & 0x7F); // flip first bit to 0 si it's ignored, then we append the other 7 bits to the integer pos++; } length = length << 7; length = length | pack[pos]; // set the leftmost bit to 1 so it's ignored, and do the same thing } pos++; FILE* customStdout = fdopen(1, "w"); FILE* source = fmemopen((void*)pack.substr(pos, length).c_str(), length, "r"); std::cout << inf(source, customStdout) << "\n";
其中inf函数基于zpipe.c实现,调用后返回Z_DATA_ERROR,错误信息为incorrect header check。我已确认pos起始位置的二进制是正确的zlib头78 9C,但问题依旧。尝试用inflateInit2(&strm, -MAX_WBITS)并跳过zlib头(pos加2)后,又返回invalid stored block lengths错误。
问题分析与修复方案
1. 对象长度解析逻辑错误
Git对象的类型和长度采用可变长度MSB编码,你的代码在处理最后一个长度字节时存在错误:
- 最后一个长度字节的最高位为0,应该仅取低7位参与长度计算,但你直接赋值了整个字节,导致长度计算偏大,取到错误的压缩数据范围。
修复后的长度解析代码:
type = (pack[pos] & 0xE0) >> 4; // 0xE0等价于11100000,语义更清晰 length = pack[pos] & 0x0F; if (pack[pos] & 0x80) { // 第一个字节最高位为1,需要继续解析长度 pos++; while (true) { length = (length << 7) | (pack[pos] & 0x7F); if (!(pack[pos] & 0x80)) { // 当前字节最高位为0,结束解析 break; } pos++; } pos++; // 移动到压缩数据起始位置 } else { pos++; // 第一个字节最高位为0,直接移动到压缩数据起始位置 }
2. 临时字符串指针失效问题
pack.substr(pos, length).c_str()返回的是临时字符串的指针,临时对象在fmemopen调用后会被销毁,导致source指向的内存无效。
修复方式:先将压缩数据保存到持久化字符串中,再传入fmemopen:
std::string compressed_data = pack.substr(pos, length); FILE* source = fmemopen((void*)compressed_data.data(), compressed_data.size(), "r");
3. Pack文件字节范围截断错误
pack.substr(20, pack.length() - 40)的硬编码跳过字节数可能不准确:HTTP响应头长度不一定是8字节,建议先打印整个pack的字节长度,核对实际的响应结构,避免截断有效数据或保留多余的HTTP响应内容。
4. 推荐:直接用zlib处理内存数据(避免FILE*开销)
跳过fmemopen的转换步骤,直接将压缩数据的内存指针传给zlib,减少中间环节的出错可能:
std::string compressed = pack.substr(pos, length); z_stream strm; strm.zalloc = Z_NULL; strm.zfree = Z_NULL; strm.opaque = Z_NULL; strm.avail_in = compressed.size(); strm.next_in = (Bytef*)compressed.data(); // 分配足够大的输出缓冲区 const size_t out_buf_size = 1024 * 1024; Bytef* out_buf = new Bytef[out_buf_size]; strm.avail_out = out_buf_size; strm.next_out = out_buf; // 初始化zlib(带zlib头用inflateInit,纯deflate数据用inflateInit2(&strm, -MAX_WBITS)) int ret = inflateInit(&strm); if (ret != Z_OK) { std::cerr << "zlib初始化失败: " << ret << "\n"; delete[] out_buf; return; } ret = inflate(&strm, Z_FINISH); if (ret == Z_STREAM_END) { std::string result((char*)out_buf, strm.total_out); std::cout << result << "\n"; } else { std::cerr << "解压错误: " << ret << ", 信息: " << strm.msg << "\n"; } inflateEnd(&strm); delete[] out_buf;
内容的提问来源于stack exchange,提问作者ljd036
相关产品推荐
相关产品推荐

