如何修改Lambda代码仅列出S3存储桶中的CSV文件?
解决AWS Lambda仅列出S3桶中CSV文件的问题
S3里没有真正的文件夹,所谓的"文件夹"其实是对象键的前缀(键名以/结尾,且Size为0),你之前用endswith('.csv')没得到预期结果,大概率是没排除这些伪文件夹条目。下面是可行的修改方案:
核心思路
先过滤掉S3中的前缀(伪文件夹),再筛选出以.csv结尾的实际文件:
- 伪文件夹的特征:
Size字段为0,且键名通常以/结尾 - 实际CSV文件:
Size大于0,且键名以.csv结尾
基础版代码(适用于文件数≤1000的情况)
import boto3 def lambda_handler(event, context): s3 = boto3.client('s3') bucket_name = 'dump' response = s3.list_objects_v2(Bucket=bucket_name) csv_files = [] if 'Contents' in response: for obj in response['Contents']: # 排除伪文件夹,只保留CSV文件 if obj['Size'] > 0 and obj['Key'].endswith('.csv'): csv_files.append(obj['Key']) print("CSV文件路径:", csv_files) return { 'statusCode': 200, 'body': csv_files }
进阶版代码(处理文件数>1000的分页场景)
S3的list_objects_v2接口一次最多返回1000个对象,如果桶内文件较多,需要循环获取所有分页数据:
import boto3 def lambda_handler(event, context): s3 = boto3.client('s3') bucket_name = 'dump' csv_files = [] continuation_token = None while True: # 根据是否有续传令牌,调用不同的list_objects_v2参数 if continuation_token: response = s3.list_objects_v2(Bucket=bucket_name, ContinuationToken=continuation_token) else: response = s3.list_objects_v2(Bucket=bucket_name) # 遍历当前页的对象,筛选CSV文件 if 'Contents' in response: for obj in response['Contents']: if obj['Size'] > 0 and obj['Key'].endswith('.csv'): csv_files.append(obj['Key']) # 判断是否还有下一页,没有则退出循环 if not response.get('IsTruncated'): break continuation_token = response['NextContinuationToken'] print("CSV文件路径:", csv_files) return { 'statusCode': 200, 'body': csv_files }
关键说明
- 必须判断
Size > 0:避免把S3的伪文件夹前缀(比如data/这类键名)误当成文件 endswith('.csv')要严格匹配:如果需要忽略大小写,可以改成obj['Key'].lower().endswith('.csv')
内容的提问来源于stack exchange,提问作者Rick
相关产品推荐
相关产品推荐

