You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python内存读写tar.gz异常:tar命令正常但Python读取为空

Python内存中读写tar.gz文件时读取为空/报错的解决方法

你遇到的核心问题是写完tar.gz到内存缓冲区后,没有将缓冲区的文件指针重置到起始位置,导致读取时从缓冲区末尾开始,自然读不到有效内容。

问题原因

当通过tarfile.open(fileobj=tar_buffer, mode="w:gz")完成归档写入后,tar_buffer的文件指针(seek位置)停在了缓冲区的末尾。此时直接以读模式打开该缓冲区,会从当前指针位置(末尾)开始读取,因此会被判定为空文件,要么返回空成员列表,要么抛出empty file异常。

修正后的完整代码

import io
import tarfile

text = "This is a test."
file_name = "test.txt"

text_buffer = io.BytesIO()
text_buffer.write(text.encode(encoding="utf-8"))

tar_buffer = io.BytesIO()

# 写入tar.gz到内存缓冲区
with tarfile.open(fileobj=tar_buffer, mode="w:gz") as archive:   
    info = tarfile.TarInfo(file_name)
    text_buffer.seek(0, io.SEEK_END)
    info.size = text_buffer.tell()
    text_buffer.seek(0, io.SEEK_SET)
    archive.addfile(info, text_buffer)

# 关键步骤:将tar_buffer的指针重置到起始位置
tar_buffer.seek(0, io.SEEK_SET)

with open("test.tar.gz", "wb") as f:
    f.write(tar_buffer.getvalue())

# 读取内存中的tar.gz
archive_contents = dict()
with tarfile.open(fileobj=tar_buffer, mode="r:*") as archive:
    for entry in archive:
        entry_fd = archive.extractfile(entry.name)
        archive_contents[entry.name] = entry_fd.read().decode("utf-8")

print(archive_contents)  # 输出: {'test.txt': 'This is a test.'}

补充说明

  • 无论写完后立即读取内存缓冲区,还是先写入磁盘再重新加载到内存,都必须确保tar_buffer的文件指针处于起始位置(执行tar_buffer.seek(0))。
  • 如果是从磁盘读取已有的tar.gz文件到内存缓冲区,同样需要在打开前执行seek(0)操作。

内容的提问来源于stack exchange,提问作者python-cat-1023

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 01:47:52