C++用lzma分块异步解压.xz文件遇格式/数据损坏错误求助
C++中LZMA分块异步解压.xz文件的问题解决
需求
在C++中使用LZMA以分块方式异步解压.xz文件,即先解压部分数据,后续再异步读取剩余内容。
当前实现
采用LZMA流解码器,维护位置变量m_pos,每次解码前调用seekg(m_pos)定位文件位置。核心代码如下:
std::size_t LzmaReader::Impl::read(void* buffer, std::size_t size){ m_inFile.open(m_path , std::ios::binary); m_strm = LZMA_STREAM_INIT; m_action = LZMA_RUN; const std::string msg = "unable to open xz file \"" + m_path + '"'; if (not m_inFile) { throw std::system_error(std::make_error_code((std::errc)errno), msg); } auto rc = lzma_stream_decoder(&m_strm, UINT64_MAX, 0); if (rc != LZMA_OK) { throw Exception(msg + ": " + Exception::getErrorCodeString(rc), rc); } m_strm.next_in = nullptr; m_strm.avail_in = 0; m_strm.next_out = m_decodedBuffer; m_strm.avail_out = BUF_SIZE; std::size_t totalBytesRead = 0; std::int64_t bytesToRead = size; m_inFile.seekg(m_pos); while(bytesToRead > 0) { if (m_strm.avail_in == 0 and !m_inFile.eof()) { m_inFile.read(reinterpret_cast<char*>(m_encodedBuffer), BUF_SIZE); // check for IO errors if (m_inFile.bad()) { const std::string msg2 = "Error while reading xz file \"" + m_path + '"'; throw std::system_error(std::make_error_code((std::errc)errno), msg); } m_strm.next_in = m_encodedBuffer; m_strm.avail_in = m_inFile.gcount(); if (m_inFile.eof()) { m_action = LZMA_FINISH; } } m_strm.next_out = static_cast<std::uint8_t *>(buffer) + totalBytesRead; m_strm.avail_out = size; rc = lzma_code(&m_strm, m_action); if (rc == LZMA_STREAM_END or m_strm.avail_out == 0) { auto writeSize = size - m_strm.avail_out; totalBytesRead += writeSize; bytesToRead -= writeSize; } if (rc != LZMA_OK) { if (rc == LZMA_STREAM_END) { break; } throw Exception("unable to read from xz file \"" + m_path + ": " + Exception::getErrorCodeString(rc),rc); } } m_pos += totalBytesRead; lzma_end(&m_strm); m_inFile.close(); return totalBytesRead; }
问题
- 第二次调用
read函数时出现LZMA_FORMAT_ERROR - 若不关闭LZMA流和文件,则出现
DATA_CORRUPT错误
问题分析
- 错误的位置定位逻辑:
m_pos记录的是解压后的字节数,但直接用seekg(m_pos)定位压缩文件的偏移完全错误——xz是压缩格式,压缩后的文件字节与解压后的字节没有线性对应关系,无法通过解压后的字节数直接定位到压缩文件的对应位置。 - 每次read重置流与文件:每次调用
read都重新打开文件、初始化LZMA流,相当于每次都从头开始解析xz文件,但又错误地seek到m_pos(解压后字节数对应的压缩文件位置),导致读取的不是合法的xz流起始,触发LZMA_FORMAT_ERROR。 - 流状态未维护:如果不关闭流和文件,上次解压的流状态未重置,继续使用会导致解码器上下文混乱,触发
DATA_CORRUPT错误。
修复方案
核心思路
- 维护LZMA流的持久状态:将流的初始化移至类的构造函数,销毁移至析构函数,避免每次
read都重置流。 - 保持文件持续打开:避免每次
read都重新打开文件,减少IO开销并维持压缩文件偏移的正确性。 - 记录压缩文件中的偏移量:而非解压后的字节数,用于后续继续读取时的定位(仅适用于支持随机访问的xz文件,即创建时指定了独立块;若为普通xz文件,需顺序解压并缓存已解压内容)。
- 启用多块支持:初始化解码器时添加
LZMA_CONCATENATED标志,支持包含多个独立xz块的文件。
修改后的代码示例
// 类构造函数:初始化流与文件 LzmaReader::Impl::Impl(const std::string& path) : m_path(path), m_compressedPos(0), m_action(LZMA_RUN) { m_inFile.open(m_path, std::ios::binary); if (!m_inFile) { const std::string msg = "unable to open xz file \"" + m_path + '"'; throw std::system_error(std::make_error_code((std::errc)errno), msg); } m_strm = LZMA_STREAM_INIT; auto rc = lzma_stream_decoder(&m_strm, UINT64_MAX, LZMA_CONCATENATED); if (rc != LZMA_OK) { throw Exception("failed to initialize lzma decoder: " + Exception::getErrorCodeString(rc), rc); } m_strm.next_in = nullptr; m_strm.avail_in = 0; } // 类析构函数:清理资源 LzmaReader::Impl::~Impl() { lzma_end(&m_strm); if (m_inFile.is_open()) { m_inFile.close(); } } std::size_t LzmaReader::Impl::read(void* buffer, std::size_t size){ if (!m_inFile.is_open()) { throw std::runtime_error("xz file is not open"); } std::size_t totalBytesRead = 0; std::uint8_t* outBuf = static_cast<std::uint8_t*>(buffer); std::size_t bytesToRead = size; // 定位到上次读取的压缩文件位置 m_inFile.seekg(m_compressedPos); while (bytesToRead > 0) { if (m_strm.avail_in == 0 && !m_inFile.eof()) { // 读取压缩数据块 m_inFile.read(reinterpret_cast<char*>(m_encodedBuffer), BUF_SIZE); if (m_inFile.bad()) { const std::string msg = "Error while reading xz file \"" + m_path + '"'; throw std::system_error(std::make_error_code((std::errc)errno), msg); } m_strm.next_in = m_encodedBuffer; m_strm.avail_in = m_inFile.gcount(); m_compressedPos += m_inFile.gcount(); // 更新压缩文件偏移 if (m_inFile.eof()) { m_action = LZMA_FINISH; } } // 设置输出缓冲区 m_strm.next_out = outBuf + totalBytesRead; m_strm.avail_out = bytesToRead; auto rc = lzma_code(&m_strm, m_action); std::size_t decodedBytes = bytesToRead - m_strm.avail_out; if (decodedBytes > 0) { totalBytesRead += decodedBytes; bytesToRead -= decodedBytes; } if (rc != LZMA_OK) { if (rc == LZMA_STREAM_END) { break; } throw Exception("unable to read from xz file \"" + m_path + ": " + Exception::getErrorCodeString(rc), rc); } } return totalBytesRead; }
注意事项
- 如果你的xz文件不是按独立块创建的(如未使用
xz --block-size=xxx参数),无法实现真正的随机分块读取,只能顺序解压并缓存已读取的内容,后续读取直接从缓存获取。 - 确保
m_encodedBuffer是类的成员变量,而非局部变量,避免每次read重新分配内存。
内容的提问来源于stack exchange,提问作者Vivek Kumar
相关产品推荐
相关产品推荐

