如何创建每小时监控S3桶的Lambda函数,仅检测当日上传且已超过1小时的文件
Fixing Your S3 Hourly Monitoring Lambda Function
Got it, let's sort out this issue for you. The core problem with your current code is that it's pulling all objects from the bucket without filtering for today's uploads, plus it's fetching the object list outside the Lambda handler (which can cause stale data due to Lambda's execution environment caching). Here's how to adjust it to only check today's files and properly identify those older than 1 hour:
Key Fixes & Explanations
- Filter for today's files first: Calculate the start of today (UTC time, since S3 uses UTC for
LastModified), and only process objects uploaded on or after this time. - Fetch object list inside the handler: Ensures you get the latest bucket contents every time the Lambda runs.
- Handle pagination: Use
list_objects_v2instead oflist_objectsto avoid missing files if your bucket has over 1000 objects. - Only send SNS when there are matching files: Avoids sending empty notifications when no old files are found.
Modified Working Code
import boto3 import os from datetime import datetime, timedelta from dateutil.tz import UTC TOPIC_ARN = 'mysnstopicarn' BUCKET_NAME = 'mybucketname' def lambda_handler(event, context): # Initialize clients inside the handler to ensure fresh connections s3_client = boto3.client('s3') sns_resource = boto3.resource('sns') sns_topic = sns_resource.Topic(TOPIC_ARN) # Calculate time boundaries (all in UTC to match S3's LastModified) now = datetime.now(UTC) one_hour_ago = now - timedelta(hours=1) # Start of today (UTC midnight) today_start = now.replace(hour=0, minute=0, second=0, microsecond=0) old_today_files = [] continuation_token = None # Handle pagination for S3 object listing (in case >1000 objects) while True: list_kwargs = { 'Bucket': BUCKET_NAME, 'Delimiter': 'Archives' } if continuation_token: list_kwargs['ContinuationToken'] = continuation_token response = s3_client.list_objects_v2(**list_kwargs) # Check if there are objects to process if 'Contents' in response: for key in response['Contents']: obj_last_modified = key['LastModified'] # First filter: only today's files if obj_last_modified >= today_start: # Second filter: older than 1 hour if obj_last_modified < one_hour_ago: # Extract just the filename (optional, adjust if you need full path) _, filename = os.path.split(key['Key']) if filename: # Skip empty filenames from split old_today_files.append(filename) # Check if there are more pages to fetch if response.get('IsTruncated'): continuation_token = response['NextContinuationToken'] else: break # Only send SNS if there are files to report if old_today_files: message = f"The following today's files are more than one hour old: {', '.join(old_today_files)}" sns_topic.publish(Message=message) print(f"Notification sent: {message}") else: print("No today's files older than 1 hour found.") return { 'statusCode': 200, 'body': f"Checked {len(old_today_files)} stale today files" }
What Changed?
- Time Boundaries: We calculate
today_startto get UTC midnight, so we only process files uploaded today. - Pagination Handling: The loop uses
list_objects_v2and checksIsTruncatedto fetch all pages of objects, so you don't miss any files. - Client Initialization: Moving
boto3client setup inside the handler ensures you don't reuse stale connections or cached data from previous Lambda executions. - Conditional SNS: We only send a notification if there are actual stale files to report, avoiding unnecessary empty messages.
- Cleaner Logic: Combined the two loops into one, filtering today's files first before checking their age, which is more efficient.
This should exactly meet your requirement: only check today's uploaded files, ignore older ones, and run hourly to flag any that have been sitting for over an hour.
内容的提问来源于stack exchange,提问作者Aadhinarayanan J
相关产品推荐
相关产品推荐

