Python读取S3中.tar内Tiff文件时os.path.exists()失效问题
问题描述
使用Python3及boto3.resource操作S3存储桶(通过s3_target.Bucket('dataset')访问),遍历桶内.tar文件并提取其中.tiff文件信息时出现异常:
编写的extract_tiff_from_tar函数中,遍历.tar内成员时,os.path.exists(member.name)仅对第一个.tiff文件返回True并打印提示,后续.tiff文件虽能被遍历并打印名称,但该判断不生效,且打开后续文件时触发FileNotFoundError。
代码示例
# Function to extract TIFF files from a tar archive def extract_tiff_from_tar(bucket, tar_key): obj = bucket.Object(tar_key) response = obj.get() # Extract .tar file from S3 object tar_buffer = io.BytesIO(response['Body'].read()) # Extract TIFF files from the .tar archive with tarfile.open(fileobj=tar_buffer, mode='r') as tar: # Determine the size of the .tar file tar_size_bytes = sum(member.size for member in tar.getmembers()) tar_size_gb = tar_size_bytes / (1024 * 1024 * 1024) print(f"Size of the tar file: {tar_size_gb:.2f} GB") for member in tar.getmembers(): # Filter the TIFF files if member.name.endswith('.tiff') and member.isfile(): print("Current file type in tar: " + str(member.name)) # Check if the file exists if os.path.exists(member.name): print('This file exists: ' + str(member.name)) # Use the try/except block to attempt to validate the existance of the TIFF try: with open(member.name, 'rb') as tiff_file: tiff = tf.imread(tiff_file) print('This is the shape of the tiff: ' + str(tiff.shape)) except FileNotFoundError as e: continue
运行输出
Size of the tar file: 1.59 GB Current file type in tar: POW-xx0 This file exists: POW-xx0 This is the shape of the tiff: (2048, 3072, 15) Current file type in tar: POW-xx1 Current file type in tar: POW-xx2
疑问:为何os.path.exists(member.name)仅对.tar内的首个文件生效,后续同属该.tar的文件无法被该方法检测到存在?
问题原因与解决方案
核心原因
你混淆了tar包内的文件成员和本地文件系统的文件:
os.path.exists(member.name)检查的是本地磁盘上是否存在同名文件,和tar包内的成员毫无关系。- 第一个文件返回True纯粹是巧合——你的本地目录刚好有一个和tar内第一个tiff同名的文件,后续文件本地不存在,所以判断失败。
- 直接用
open(member.name)尝试打开文件,本质是在本地查找该文件,而非从tar包内读取内容,因此后续文件会触发FileNotFoundError。
正确解决方案
要读取tar包内的tiff文件,不需要提取到本地,直接从tar包的内存流中读取即可:
修改后的代码如下:
import io import tarfile import tensorflow as tf def extract_tiff_from_tar(bucket, tar_key): obj = bucket.Object(tar_key) response = obj.get() tar_buffer = io.BytesIO(response['Body'].read()) with tarfile.open(fileobj=tar_buffer, mode='r') as tar: tar_size_bytes = sum(member.size for member in tar.getmembers()) tar_size_gb = tar_size_bytes / (1024 * 1024 * 1024) print(f"Size of the tar file: {tar_size_gb:.2f} GB") for member in tar.getmembers(): if member.name.endswith('.tiff') and member.isfile(): print(f"Processing TIFF file: {member.name}") # 直接从tar包中读取文件内容,无需本地文件 with tar.extractfile(member) as tiff_file: tiff = tf.imread(tiff_file) print(f"Shape of {member.name}: {tiff.shape}")
关键修改点
- 移除无意义的
os.path.exists判断,因为它和tar包内的文件无关。 - 用
tar.extractfile(member)替代open(member.name),直接从tar包的内存流中读取指定成员的内容,无需将文件提取到本地磁盘。 - 移除了多余的try/except块(若需异常处理可保留,但核心逻辑是从tar内读取)。
内容的提问来源于stack exchange,提问作者tds
相关产品推荐
相关产品推荐

