You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中解压.tar.zst文件的高效方法及代码报错排查

.tar.zst打包与解压的Python代码问题修正

问题

我正在使用zstandard库进行.tar.zst文件的打包与解压,试图在Python中复现以下bash命令:

curl -O - http://download.com/file.tar.zst | zstd -d - | tar xf -

看到相关讨论提到Python的tarfile模块支持任意流,使用zstandard应该很简单,但在使用流时遇到了问题。

我尝试编写了如下代码:

import requests
import zstandard
import tarfile

# Write .tar.zst archive
path = "/some/path/to/tar"
to_zstd = zstd.ZstdCompressor()
with tempfile.NamedTemporaryFile() as temp_tar_file:
  with tarfile.open(name="file.tar.zst", fileobj=to_zstd.stream_writer(temp_tar_file), mode="w|") as tf:
    tf.add(path)

# Upload generated .tar.zst

# Extract .tar.zst
with closing(requests.get(url, stream=True) as r:
  tardir = tempfile.mkdtemp()
  unzstd = zstd.ZstdDecompressor()
  with io.BytesIO() as tarbyte:
    unzstd.copy_stream(resp.raw, tarbyte)
    with tarfile.open(fileobj=tarbyte, mode="r|") as tf:
      tf.extractall(path=tardir)

运行代码时出现tarfile.ReadError("empty file"),使用zstd -d file.tar.zst命令解压时则提示Read error (39) : premature end。该文件大小非零且随压缩内容变化,但压缩过程显然存在问题,请问我哪里出错了?

问题分析与修正

核心错误点

  • 压缩流程:zstd流写入器未正确完成收尾,tarfile关闭后缓冲区数据未刷新到文件,导致压缩文件不完整。
  • 解压流程:存在语法错误(closing使用不当)、变量名错误(resp.raw应为r.raw)、BytesIO指针未重置到开头,导致tarfile读取空内容。

修正后的压缩代码

import tempfile
import zstandard
import tarfile

path = "/some/path/to/tar"
to_zstd = zstd.ZstdCompressor()

# 保留临时文件以便后续上传,delete=False避免自动删除
with tempfile.NamedTemporaryFile(delete=False) as temp_tar_file:
    # 先创建zstd流写入器,确保tar流写完后zstd能完成压缩收尾
    with to_zstd.stream_writer(temp_tar_file) as zstd_stream:
        with tarfile.open(fileobj=zstd_stream, mode="w|") as tf:
            tf.add(path)

# 此时temp_tar_file.name即为完整的.tar.zst文件路径,可用于上传
# 上传完成后记得手动删除临时文件:import os; os.unlink(temp_tar_file.name)

修正后的解压代码

import requests
import zstandard
import tarfile
from contextlib import closing
import tempfile

url = "http://download.com/file.tar.zst"
with closing(requests.get(url, stream=True)) as r:
    r.raise_for_status()  # 确保请求成功,避免读取错误响应
    tardir = tempfile.mkdtemp()
    unzstd = zstd.ZstdDecompressor()
    
    # 直接将解压流传给tarfile,无需中转BytesIO,节省内存且更高效
    with unzstd.stream_reader(r.raw) as zstd_stream:
        with tarfile.open(fileobj=zstd_stream, mode="r|") as tf:
            tf.extractall(path=tardir)

# 若需使用BytesIO中转,需添加seek(0)重置指针:
# import io
# with io.BytesIO() as tarbyte:
#     unzstd.copy_stream(r.raw, tarbyte)
#     tarbyte.seek(0)  # 将文件指针移到开头,否则tarfile从末尾读取
#     with tarfile.open(fileobj=tarbyte, mode="r|") as tf:
#         tf.extractall(path=tardir)

关键说明

  • 压缩时,将zstd流写入器的with块放在tarfile外层,确保tar数据全部写入后,zstd流能自动刷新缓冲区并写入压缩结束标记,彻底解决"premature end"错误。
  • 解压时,直接流式处理响应内容,避免将整个文件加载到内存,同时修正语法和变量错误,确保tarfile能读取到完整的解压数据。

内容的提问来源于stack exchange,提问作者drowsily.throbbing872

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 13:13:14