You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C++用lzma分块异步解压.xz文件遇格式/数据损坏错误求助

C++中LZMA分块异步解压.xz文件的问题解决

需求

在C++中使用LZMA以分块方式异步解压.xz文件,即先解压部分数据,后续再异步读取剩余内容。

当前实现

采用LZMA流解码器,维护位置变量m_pos,每次解码前调用seekg(m_pos)定位文件位置。核心代码如下:

std::size_t LzmaReader::Impl::read(void* buffer, std::size_t size){
    m_inFile.open(m_path , std::ios::binary);
    m_strm = LZMA_STREAM_INIT;
    m_action = LZMA_RUN;
    const std::string msg = "unable to open xz file \"" + m_path + '"';

    if (not m_inFile) {
        throw std::system_error(std::make_error_code((std::errc)errno), msg);
    }
    auto rc = lzma_stream_decoder(&m_strm, UINT64_MAX, 0);
    if (rc != LZMA_OK) {
        throw Exception(msg + ": " + Exception::getErrorCodeString(rc), rc);
    }

    m_strm.next_in = nullptr;
    m_strm.avail_in = 0;
    m_strm.next_out = m_decodedBuffer;
    m_strm.avail_out = BUF_SIZE;

    std::size_t totalBytesRead = 0;

    std::int64_t bytesToRead = size;
    m_inFile.seekg(m_pos);
    while(bytesToRead > 0)
    {
        if (m_strm.avail_in == 0 and !m_inFile.eof())
        {
            m_inFile.read(reinterpret_cast<char*>(m_encodedBuffer), BUF_SIZE);

            // check for IO errors
            if (m_inFile.bad())
            {
                const std::string msg2 = "Error while reading xz file \"" + m_path + '"';
                throw std::system_error(std::make_error_code((std::errc)errno), msg);
            }

            m_strm.next_in = m_encodedBuffer;
            m_strm.avail_in = m_inFile.gcount();

            if (m_inFile.eof())
            {
                m_action = LZMA_FINISH;
            }
        }

        m_strm.next_out = static_cast<std::uint8_t *>(buffer) + totalBytesRead;
        m_strm.avail_out = size;

        rc = lzma_code(&m_strm, m_action);

        if (rc == LZMA_STREAM_END or m_strm.avail_out == 0)
        {
            auto writeSize = size - m_strm.avail_out;
            totalBytesRead += writeSize;
            bytesToRead -= writeSize;
        }

        if (rc != LZMA_OK)
        {
            if (rc == LZMA_STREAM_END)
            {
                break;
            }

            throw Exception("unable to read from xz file \"" + m_path + ": "
                                + Exception::getErrorCodeString(rc),rc);
        }
    }

    m_pos += totalBytesRead;
    lzma_end(&m_strm);
    m_inFile.close();
    return totalBytesRead;
}

问题

  • 第二次调用read函数时出现LZMA_FORMAT_ERROR
  • 若不关闭LZMA流和文件,则出现DATA_CORRUPT错误

问题分析

  1. 错误的位置定位逻辑:m_pos记录的是解压后的字节数,但直接用seekg(m_pos)定位压缩文件的偏移完全错误——xz是压缩格式,压缩后的文件字节与解压后的字节没有线性对应关系,无法通过解压后的字节数直接定位到压缩文件的对应位置。
  2. 每次read重置流与文件:每次调用read都重新打开文件、初始化LZMA流,相当于每次都从头开始解析xz文件,但又错误地seek到m_pos(解压后字节数对应的压缩文件位置),导致读取的不是合法的xz流起始,触发LZMA_FORMAT_ERROR。
  3. 流状态未维护:如果不关闭流和文件,上次解压的流状态未重置,继续使用会导致解码器上下文混乱,触发DATA_CORRUPT错误。

修复方案

核心思路

  • 维护LZMA流的持久状态:将流的初始化移至类的构造函数,销毁移至析构函数,避免每次read都重置流。
  • 保持文件持续打开:避免每次read都重新打开文件,减少IO开销并维持压缩文件偏移的正确性。
  • 记录压缩文件中的偏移量:而非解压后的字节数,用于后续继续读取时的定位(仅适用于支持随机访问的xz文件,即创建时指定了独立块;若为普通xz文件,需顺序解压并缓存已解压内容)。
  • 启用多块支持:初始化解码器时添加LZMA_CONCATENATED标志,支持包含多个独立xz块的文件。

修改后的代码示例

// 类构造函数:初始化流与文件
LzmaReader::Impl::Impl(const std::string& path) 
    : m_path(path), m_compressedPos(0), m_action(LZMA_RUN) {
    m_inFile.open(m_path, std::ios::binary);
    if (!m_inFile) {
        const std::string msg = "unable to open xz file \"" + m_path + '"';
        throw std::system_error(std::make_error_code((std::errc)errno), msg);
    }

    m_strm = LZMA_STREAM_INIT;
    auto rc = lzma_stream_decoder(&m_strm, UINT64_MAX, LZMA_CONCATENATED);
    if (rc != LZMA_OK) {
        throw Exception("failed to initialize lzma decoder: " + Exception::getErrorCodeString(rc), rc);
    }

    m_strm.next_in = nullptr;
    m_strm.avail_in = 0;
}

// 类析构函数:清理资源
LzmaReader::Impl::~Impl() {
    lzma_end(&m_strm);
    if (m_inFile.is_open()) {
        m_inFile.close();
    }
}

std::size_t LzmaReader::Impl::read(void* buffer, std::size_t size){
    if (!m_inFile.is_open()) {
        throw std::runtime_error("xz file is not open");
    }

    std::size_t totalBytesRead = 0;
    std::uint8_t* outBuf = static_cast<std::uint8_t*>(buffer);
    std::size_t bytesToRead = size;

    // 定位到上次读取的压缩文件位置
    m_inFile.seekg(m_compressedPos);

    while (bytesToRead > 0) {
        if (m_strm.avail_in == 0 && !m_inFile.eof()) {
            // 读取压缩数据块
            m_inFile.read(reinterpret_cast<char*>(m_encodedBuffer), BUF_SIZE);
            if (m_inFile.bad()) {
                const std::string msg = "Error while reading xz file \"" + m_path + '"';
                throw std::system_error(std::make_error_code((std::errc)errno), msg);
            }

            m_strm.next_in = m_encodedBuffer;
            m_strm.avail_in = m_inFile.gcount();
            m_compressedPos += m_inFile.gcount(); // 更新压缩文件偏移

            if (m_inFile.eof()) {
                m_action = LZMA_FINISH;
            }
        }

        // 设置输出缓冲区
        m_strm.next_out = outBuf + totalBytesRead;
        m_strm.avail_out = bytesToRead;

        auto rc = lzma_code(&m_strm, m_action);

        std::size_t decodedBytes = bytesToRead - m_strm.avail_out;
        if (decodedBytes > 0) {
            totalBytesRead += decodedBytes;
            bytesToRead -= decodedBytes;
        }

        if (rc != LZMA_OK) {
            if (rc == LZMA_STREAM_END) {
                break;
            }
            throw Exception("unable to read from xz file \"" + m_path + ": "
                            + Exception::getErrorCodeString(rc), rc);
        }
    }

    return totalBytesRead;
}

注意事项

  • 如果你的xz文件不是按独立块创建的(如未使用xz --block-size=xxx参数),无法实现真正的随机分块读取,只能顺序解压并缓存已读取的内容,后续读取直接从缓存获取。
  • 确保m_encodedBuffer是类的成员变量,而非局部变量,避免每次read重新分配内存。

内容的提问来源于stack exchange,提问作者Vivek Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 11:19:50