You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取S3中.tar内Tiff文件时os.path.exists()失效问题

问题描述

使用Python3及boto3.resource操作S3存储桶(通过s3_target.Bucket('dataset')访问),遍历桶内.tar文件并提取其中.tiff文件信息时出现异常:
编写的extract_tiff_from_tar函数中,遍历.tar内成员时,os.path.exists(member.name)仅对第一个.tiff文件返回True并打印提示,后续.tiff文件虽能被遍历并打印名称,但该判断不生效,且打开后续文件时触发FileNotFoundError。

代码示例

# Function to extract TIFF files from a tar archive
def extract_tiff_from_tar(bucket, tar_key):
    obj = bucket.Object(tar_key)
    response = obj.get()
    
    # Extract .tar file from S3 object
    tar_buffer = io.BytesIO(response['Body'].read())

    # Extract TIFF files from the .tar archive
    with tarfile.open(fileobj=tar_buffer, mode='r') as tar:

        # Determine the size of the .tar file
        tar_size_bytes = sum(member.size for member in tar.getmembers())
        tar_size_gb = tar_size_bytes / (1024 * 1024 * 1024)
        print(f"Size of the tar file: {tar_size_gb:.2f} GB")

        for member in tar.getmembers():
            # Filter the TIFF files
            if member.name.endswith('.tiff') and member.isfile():
                print("Current file type in tar: " + str(member.name))
                # Check if the file exists
                if os.path.exists(member.name):
                    print('This file exists: ' + str(member.name))

                # Use the try/except block to attempt to validate the existance of the TIFF
                try:
                    with open(member.name, 'rb') as tiff_file:
                        tiff = tf.imread(tiff_file)
                        print('This is the shape of the tiff: ' + str(tiff.shape))

                except FileNotFoundError as e:
                    continue

运行输出

Size of the tar file: 1.59 GB
Current file type in tar: POW-xx0
This file exists: POW-xx0
This is the shape of the tiff: (2048, 3072, 15)
Current file type in tar: POW-xx1
Current file type in tar: POW-xx2

疑问:为何os.path.exists(member.name)仅对.tar内的首个文件生效,后续同属该.tar的文件无法被该方法检测到存在?


问题原因与解决方案

核心原因

你混淆了tar包内的文件成员和本地文件系统的文件:

  • os.path.exists(member.name)检查的是本地磁盘上是否存在同名文件,和tar包内的成员毫无关系。
  • 第一个文件返回True纯粹是巧合——你的本地目录刚好有一个和tar内第一个tiff同名的文件,后续文件本地不存在,所以判断失败。
  • 直接用open(member.name)尝试打开文件,本质是在本地查找该文件,而非从tar包内读取内容,因此后续文件会触发FileNotFoundError。

正确解决方案

要读取tar包内的tiff文件,不需要提取到本地,直接从tar包的内存流中读取即可:

修改后的代码如下:

import io
import tarfile
import tensorflow as tf

def extract_tiff_from_tar(bucket, tar_key):
    obj = bucket.Object(tar_key)
    response = obj.get()
    
    tar_buffer = io.BytesIO(response['Body'].read())

    with tarfile.open(fileobj=tar_buffer, mode='r') as tar:
        tar_size_bytes = sum(member.size for member in tar.getmembers())
        tar_size_gb = tar_size_bytes / (1024 * 1024 * 1024)
        print(f"Size of the tar file: {tar_size_gb:.2f} GB")

        for member in tar.getmembers():
            if member.name.endswith('.tiff') and member.isfile():
                print(f"Processing TIFF file: {member.name}")
                # 直接从tar包中读取文件内容,无需本地文件
                with tar.extractfile(member) as tiff_file:
                    tiff = tf.imread(tiff_file)
                    print(f"Shape of {member.name}: {tiff.shape}")

关键修改点

  • 移除无意义的os.path.exists判断,因为它和tar包内的文件无关。
  • 用tar.extractfile(member)替代open(member.name),直接从tar包的内存流中读取指定成员的内容,无需将文件提取到本地磁盘。
  • 移除了多余的try/except块(若需异常处理可保留,但核心逻辑是从tar内读取)。

内容的提问来源于stack exchange,提问作者tds

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 20:35:35