如何在Lambda函数中指定不解压的tar文件以避免超时?
如何在Lambda的untar操作中排除指定文件
当然可以,你只需要在遍历tar包内文件的逻辑里添加过滤规则,跳过不需要解压的文件,就能减少Lambda的执行操作量,避免超时。
修改思路
- 先定义需要排除的文件规则(精确文件名、后缀、前缀都可以)
- 遍历tar包成员时,判断当前文件是否符合排除规则,符合则跳过,否则执行解压上传
修改后的代码
def untar_file(zip_key, source_bucket, source_path, file): # 定义需要排除的文件列表,按需修改 exclude_files = ["log.txt", "temp.dat"] # 也可设置后缀过滤,比如排除所有临时文件:exclude_suffix = ".tmp" zip_obj = s3_resource.Object(bucket_name=source_bucket, key=zip_key) buffer = BytesIO(zip_obj.get()["Body"].read()) with tarfile.open(fileobj=buffer, mode=('r:gz')) as z: for member in z.getmembers(): # 跳过目录(如果不需要处理目录的话) if member.isdir(): continue # 跳过排除列表中的文件 if member.name in exclude_files: continue # 若用后缀过滤,替换为:if member.name.endswith(exclude_suffix): continue # 若用前缀过滤,替换为:if member.name.startswith("temp_"): continue # 仅处理需要保留的文件 s3_resource.meta.client.upload_fileobj( z.extractfile(member), Bucket=source_bucket, Key=f"{source_path}/{d1}/{member.name}.csv" # 确保d1变量已正确定义 ) copy_objects(zip_key, source_bucket, source_path, file)
额外说明
- 过滤规则可以根据实际需求调整,比如排除特定目录、匹配正则表达式等
- 确认
d1变量在函数内已正确初始化,否则会引发NameError - 减少不必要的文件处理后,Lambda的执行时间会明显缩短,超时概率大幅降低
内容的提问来源于stack exchange,提问作者facepalmdev7
相关产品推荐
相关产品推荐

