Python中解压.tar.zst文件的高效方法及代码报错排查
.tar.zst打包与解压的Python代码问题修正
问题
我正在使用zstandard库进行.tar.zst文件的打包与解压,试图在Python中复现以下bash命令:
curl -O - http://download.com/file.tar.zst | zstd -d - | tar xf -
看到相关讨论提到Python的tarfile模块支持任意流,使用zstandard应该很简单,但在使用流时遇到了问题。
我尝试编写了如下代码:
import requests import zstandard import tarfile # Write .tar.zst archive path = "/some/path/to/tar" to_zstd = zstd.ZstdCompressor() with tempfile.NamedTemporaryFile() as temp_tar_file: with tarfile.open(name="file.tar.zst", fileobj=to_zstd.stream_writer(temp_tar_file), mode="w|") as tf: tf.add(path) # Upload generated .tar.zst # Extract .tar.zst with closing(requests.get(url, stream=True) as r: tardir = tempfile.mkdtemp() unzstd = zstd.ZstdDecompressor() with io.BytesIO() as tarbyte: unzstd.copy_stream(resp.raw, tarbyte) with tarfile.open(fileobj=tarbyte, mode="r|") as tf: tf.extractall(path=tardir)
运行代码时出现tarfile.ReadError("empty file"),使用zstd -d file.tar.zst命令解压时则提示Read error (39) : premature end。该文件大小非零且随压缩内容变化,但压缩过程显然存在问题,请问我哪里出错了?
问题分析与修正
核心错误点
- 压缩流程:zstd流写入器未正确完成收尾,tarfile关闭后缓冲区数据未刷新到文件,导致压缩文件不完整。
- 解压流程:存在语法错误(
closing使用不当)、变量名错误(resp.raw应为r.raw)、BytesIO指针未重置到开头,导致tarfile读取空内容。
修正后的压缩代码
import tempfile import zstandard import tarfile path = "/some/path/to/tar" to_zstd = zstd.ZstdCompressor() # 保留临时文件以便后续上传,delete=False避免自动删除 with tempfile.NamedTemporaryFile(delete=False) as temp_tar_file: # 先创建zstd流写入器,确保tar流写完后zstd能完成压缩收尾 with to_zstd.stream_writer(temp_tar_file) as zstd_stream: with tarfile.open(fileobj=zstd_stream, mode="w|") as tf: tf.add(path) # 此时temp_tar_file.name即为完整的.tar.zst文件路径,可用于上传 # 上传完成后记得手动删除临时文件:import os; os.unlink(temp_tar_file.name)
修正后的解压代码
import requests import zstandard import tarfile from contextlib import closing import tempfile url = "http://download.com/file.tar.zst" with closing(requests.get(url, stream=True)) as r: r.raise_for_status() # 确保请求成功,避免读取错误响应 tardir = tempfile.mkdtemp() unzstd = zstd.ZstdDecompressor() # 直接将解压流传给tarfile,无需中转BytesIO,节省内存且更高效 with unzstd.stream_reader(r.raw) as zstd_stream: with tarfile.open(fileobj=zstd_stream, mode="r|") as tf: tf.extractall(path=tardir) # 若需使用BytesIO中转,需添加seek(0)重置指针: # import io # with io.BytesIO() as tarbyte: # unzstd.copy_stream(r.raw, tarbyte) # tarbyte.seek(0) # 将文件指针移到开头,否则tarfile从末尾读取 # with tarfile.open(fileobj=tarbyte, mode="r|") as tf: # tf.extractall(path=tardir)
关键说明
- 压缩时,将zstd流写入器的
with块放在tarfile外层,确保tar数据全部写入后,zstd流能自动刷新缓冲区并写入压缩结束标记,彻底解决"premature end"错误。 - 解压时,直接流式处理响应内容,避免将整个文件加载到内存,同时修正语法和变量错误,确保tarfile能读取到完整的解压数据。
内容的提问来源于stack exchange,提问作者drowsily.throbbing872
相关产品推荐
相关产品推荐

