You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 3.6.4使用bz2.BZ2Decompressor解压维基数据文件失败求助

解决Python的bz2.BZ2Decompressor解压维基媒体bz2文件失败的问题

我之前处理维基媒体的大体积bz2数据时也碰到过类似的坑——7z能正常解压,但用bz2.BZ2Decompressor就卡第一块输出为空,折腾了好久才搞明白原因,给你分享下解决方案:

问题根源

维基媒体的这个bz2文件是多流压缩格式的,而Python的bz2.BZ2Decompressor默认是针对单流压缩数据设计的。当它遇到多流数据时,第一块解压后可能返回空(因为它在等待后续的流数据),如果你的代码没处理这种多流场景,就会直接判定解压失败。

靠谱的解决方案

方案1:用高层API bz2.open()(推荐)

这个方法会自动处理多流压缩的情况,比手动用BZ2Decompressor省心太多,代码也简洁:

import bz2

# 直接打开压缩文件并逐块解压写入输出
with bz2.open('enwiktionary-latest-pages-meta-current.xml.bz2', 'rb') as compressed_file:
    with open('decompressed_output.xml', 'wb') as output_file:
        # 按1MB块读取解压,避免内存占用过高
        for chunk in iter(lambda: compressed_file.read(1024 * 1024), b''):
            output_file.write(chunk)

方案2:手动处理多流(如果必须用BZ2Decompressor)

如果你的业务逻辑必须用BZ2Decompressor,那得在代码里加上多流处理的逻辑,循环重置解压器直到所有流都处理完:

import bz2
from queue import Queue

def decompression(qin: Queue, qout: Queue):
    decompressor = bz2.BZ2Decompressor()
    while not qin.empty():
        data = qin.get()
        try:
            # 解压当前块数据
            decompressed_data = decompressor.decompress(data)
            if decompressed_data:
                qout.put(decompressed_data)
            
            # 检查是否处理完当前流,还有后续流的话重置解压器
            while decompressor.eof:
                decompressor = bz2.BZ2Decompressor()
                # 处理上一个流剩余的未使用数据
                if decompressor.unused_data:
                    decompressed_data = decompressor.decompress(decompressor.unused_data)
                    if decompressed_data:
                        qout.put(decompressed_data)
        except Exception as e:
            print(f"解压过程出错: {str(e)}")
            break
    # 处理最后可能的剩余数据
    try:
        final_data = decompressor.flush()
        if final_data:
            qout.put(final_data)
    except Exception as e:
        print(f"收尾解压出错: {str(e)}")

# 示例:用生成器分块读取文件作为输入队列的数据源
def file_chunk_generator(file_path, chunk_size=1024*1024):
    with open(file_path, 'rb') as f:
        while chunk := f.read(chunk_size):
            yield chunk

# 使用示例
input_queue = Queue()
output_queue = Queue()
for chunk in file_chunk_generator('enwiktionary-latest-pages-meta-current.xml.bz2'):
    input_queue.put(chunk)

decompression(input_queue, output_queue)

# 从输出队列读取解压后的数据写入文件
with open('output.xml', 'wb') as f:
    while not output_queue.empty():
        f.write(output_queue.get())

为什么7z能正常工作?

7z的bz2解压实现完全支持多流压缩格式,会自动识别并处理多个压缩流,不需要用户额外操作,所以能顺利解开文件。

内容的提问来源于stack exchange,提问作者Jonathan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:11:09