如何用Lambda+Boto3获取S3指定文件夹最新CSV文件及解决权限报错
解决方案:修复权限问题并实现S3指定文件夹最新CSV文件获取
先给你拆解问题,一步步解决:
1. 解决AllAccessDisabled权限错误
这个错误是因为你的Lambda执行角色没有足够的S3操作权限。你需要给Lambda绑定的IAM角色添加以下权限策略:
{ "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "s3:ListObjectsV2", "s3:GetObject" ], "Resource": [ "arn:aws:s3:::BUCKET", "arn:aws:s3:::BUCKET/Output/*" ] } ] }
- 把
BUCKET替换成你的实际存储桶名称 - 这里遵循最小权限原则,只授予了Output文件夹的访问权限;如果需要访问Input文件夹,可以添加对应的
Resource条目
2. 修正代码:指定文件夹+按Last Modified取最新文件
你的原代码存在两个关键问题:没有指定目标文件夹、排序后取到的是最早而非最新的文件,同时缺少异常处理逻辑。以下是修正后的完整代码:
import json import boto3 def lambda_handler(event, context): s3_client = boto3.client('s3') bucket_name = 'BUCKET' target_folder = 'Output/' # 指定要搜索的文件夹,结尾的/必须保留 try: # 列出指定文件夹下的所有对象 response = s3_client.list_objects_v2(Bucket=bucket_name, Prefix=target_folder) # 检查文件夹是否为空 if 'Contents' not in response: return { "statusCode": 404, "body": json.dumps("No files found in the specified folder") } # 过滤掉文件夹本身的条目,只保留文件 file_objects = [obj for obj in response['Contents'] if not obj['Key'].endswith('/')] # 按Last Modified降序排序,取第一个即为最新文件 get_last_modified = lambda obj: obj['LastModified'] latest_file = sorted(file_objects, key=get_last_modified, reverse=True)[0] latest_file_key = latest_file['Key'] # 生成有效期2小时的预签名URL presigned_url = s3_client.generate_presigned_url( 'get_object', Params={'Bucket': bucket_name, 'Key': latest_file_key}, ExpiresIn=7200, HttpMethod='GET' ) # 返回303重定向,方便PowerBI/Tableau直接获取文件 return { "statusCode": 303, "headers": {'Location': presigned_url} } except Exception as e: return { "statusCode": 500, "body": json.dumps(f"Error occurred: {str(e)}") }
3. 测试验证
- 将代码中的
BUCKET替换为你的实际存储桶名称 - 确认Lambda的IAM角色已添加正确的权限策略
- 触发Lambda函数,返回的
Location响应头即为Output文件夹中最新CSV文件的预签名URL - 复制该URL到浏览器测试,确认能正常下载文件后,即可在PowerBI/Tableau中使用API Gateway的地址自动获取最新数据
内容的提问来源于stack exchange,提问作者Filip van der Pol
相关产品推荐
相关产品推荐

