如何在S3文件更新后自动批量失效CloudFront缓存?
Absolutely! This is a perfect fit for AWS Lambda—you can absolutely automate full CloudFront cache invalidations (using *) right after your S3 bucket gets updated via FTP. Let me break down exactly how to set this up, including cost-saving tips to avoid unnecessary charges:
First, you need to trigger your Lambda function whenever files are added or overwritten in your S3 bucket. Since you’re replacing nearly all files, use the s3:ObjectCreated:* event type (it covers PUT, POST, COPY, and other write operations):
- Navigate to your S3 bucket > Properties > Event notifications > Create event
- Name your event (e.g.,
FTP-S3-Update-Trigger) - Under Event types, select all options under
ObjectCreated - For Destination, choose
Lambda functionand select (or create) your target Lambda function - Critical note: Ensure your Lambda execution role has permissions to read S3 event data and run CloudFront invalidations (more on this later)
- Name your event (e.g.,
Here’s a straightforward Python function that triggers a full cache invalidation using /* (equivalent to * for full distribution coverage):
import boto3 def lambda_handler(event, context): cloudfront_client = boto3.client('cloudfront') # Replace with your actual CloudFront distribution ID DISTRIBUTION_ID = 'YOUR_DISTRO_ID' try: # Initiate full cache invalidation invalidation_response = cloudfront_client.create_invalidation( DistributionId=DISTRIBUTION_ID, InvalidationBatch={ 'Paths': { 'Quantity': 1, 'Items': ['/*'] }, 'CallerReference': str(context.request_id) } ) print(f"Full invalidation started successfully. ID: {invalidation_response['Invalidation']['Id']}") return { 'statusCode': 200, 'body': f"Invalidation initiated: {invalidation_response['Invalidation']['Id']}" } except Exception as e: print(f"Error creating invalidation: {str(e)}") raise e
- Key details:
['/*']is the correct syntax for a full distribution invalidation (counts as just 1 path for billing)CallerReferenceuses Lambda’s unique request ID to ensure each invalidation request is distinct (required by CloudFront)
Since you’re overwriting dozens/hundreds of files, S3 will trigger Lambda once per file—leading to redundant invalidation requests (and wasted cost). Fix this with a batching/deduplication layer:
- Option 1: SQS Delay Queue + DynamoDB State Tracking
- Route S3 events to an SQS queue with a 5-10 minute delay
- Configure Lambda to poll this queue, and check DynamoDB if an invalidation was already initiated in the last 15 minutes. If yes, skip the request; if not, run the invalidation and update DynamoDB with the timestamp.
- Option 2: EventBridge Scheduler
- Send S3 events to EventBridge, then create a scheduled rule that runs every 10 minutes. The rule triggers Lambda only if there were S3 update events in that window. This ensures you only run one invalidation per batch, no matter how many files were updated.
Your choice to use full invalidation (/*) is exactly the right move for cost control:
- AWS charges CloudFront invalidations by the number of paths. A full invalidation counts as 1 path, whereas invalidating 1000 individual files counts as 1000 paths. This will drastically reduce your costs compared to per-file invalidations.
- Lambda costs are negligible here—even with occasional runs, execution time is just a few seconds, well within the free tier limits for most use cases.
Attach this policy to your Lambda execution role to grant necessary access:
{ "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": "cloudfront:CreateInvalidation", "Resource": "arn:aws:cloudfront::YOUR_AWS_ACCOUNT_ID:distribution/YOUR_DISTRO_ID" }, { "Effect": "Allow", "Action": [ "logs:CreateLogGroup", "logs:CreateLogStream", "logs:PutLogEvents" ], "Resource": "arn:aws:logs:*:*:*" } ] }
内容的提问来源于stack exchange,提问作者AJ222

