You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何创建每小时监控S3桶的Lambda函数,仅检测当日上传且已超过1小时的文件

Fixing Your S3 Hourly Monitoring Lambda Function

Got it, let's sort out this issue for you. The core problem with your current code is that it's pulling all objects from the bucket without filtering for today's uploads, plus it's fetching the object list outside the Lambda handler (which can cause stale data due to Lambda's execution environment caching). Here's how to adjust it to only check today's files and properly identify those older than 1 hour:

Key Fixes & Explanations

  • Filter for today's files first: Calculate the start of today (UTC time, since S3 uses UTC for LastModified), and only process objects uploaded on or after this time.
  • Fetch object list inside the handler: Ensures you get the latest bucket contents every time the Lambda runs.
  • Handle pagination: Use list_objects_v2 instead of list_objects to avoid missing files if your bucket has over 1000 objects.
  • Only send SNS when there are matching files: Avoids sending empty notifications when no old files are found.

Modified Working Code

import boto3
import os
from datetime import datetime, timedelta
from dateutil.tz import UTC

TOPIC_ARN = 'mysnstopicarn'
BUCKET_NAME = 'mybucketname'

def lambda_handler(event, context):
    # Initialize clients inside the handler to ensure fresh connections
    s3_client = boto3.client('s3')
    sns_resource = boto3.resource('sns')
    sns_topic = sns_resource.Topic(TOPIC_ARN)
    
    # Calculate time boundaries (all in UTC to match S3's LastModified)
    now = datetime.now(UTC)
    one_hour_ago = now - timedelta(hours=1)
    # Start of today (UTC midnight)
    today_start = now.replace(hour=0, minute=0, second=0, microsecond=0)
    
    old_today_files = []
    continuation_token = None
    
    # Handle pagination for S3 object listing (in case >1000 objects)
    while True:
        list_kwargs = {
            'Bucket': BUCKET_NAME,
            'Delimiter': 'Archives'
        }
        if continuation_token:
            list_kwargs['ContinuationToken'] = continuation_token
            
        response = s3_client.list_objects_v2(**list_kwargs)
        # Check if there are objects to process
        if 'Contents' in response:
            for key in response['Contents']:
                obj_last_modified = key['LastModified']
                # First filter: only today's files
                if obj_last_modified >= today_start:
                    # Second filter: older than 1 hour
                    if obj_last_modified < one_hour_ago:
                        # Extract just the filename (optional, adjust if you need full path)
                        _, filename = os.path.split(key['Key'])
                        if filename:  # Skip empty filenames from split
                            old_today_files.append(filename)
        
        # Check if there are more pages to fetch
        if response.get('IsTruncated'):
            continuation_token = response['NextContinuationToken']
        else:
            break
    
    # Only send SNS if there are files to report
    if old_today_files:
        message = f"The following today's files are more than one hour old: {', '.join(old_today_files)}"
        sns_topic.publish(Message=message)
        print(f"Notification sent: {message}")
    else:
        print("No today's files older than 1 hour found.")
    
    return {
        'statusCode': 200,
        'body': f"Checked {len(old_today_files)} stale today files"
    }

What Changed?

  1. Time Boundaries: We calculate today_start to get UTC midnight, so we only process files uploaded today.
  2. Pagination Handling: The loop uses list_objects_v2 and checks IsTruncated to fetch all pages of objects, so you don't miss any files.
  3. Client Initialization: Moving boto3 client setup inside the handler ensures you don't reuse stale connections or cached data from previous Lambda executions.
  4. Conditional SNS: We only send a notification if there are actual stale files to report, avoiding unnecessary empty messages.
  5. Cleaner Logic: Combined the two loops into one, filtering today's files first before checking their age, which is more efficient.

This should exactly meet your requirement: only check today's uploaded files, ignore older ones, and run hourly to flag any that have been sitting for over an hour.

内容的提问来源于stack exchange,提问作者Aadhinarayanan J

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.01 00:14:05