AWS S3到SQS到Lambda架构下大数量对象创建时事件丢失问题排查与配置优化咨询
First, let's break down where the problem might be and how to fix it based on your setup:
Why are events going missing?
From your description, 10% of events aren't even reaching SQS when you scale up to 1000+ objects. Since your Lambda logs never see these events, the issue starts at the S3 -> SQS step, not the SQS -> Lambda handoff. Let's tackle that first, then optimize the SQS/Lambda side to ensure no events slip through once they hit SQS.
1. Fix S3 to SQS Event Delivery
Add a Dead-Letter Queue (DLQ) for S3 Event Notifications
S3 event notifications use an at-least-once delivery model, but if SQS can't accept messages (e.g., temporary throttling, permission blips), S3 will retry a few times before dropping the event. To catch these lost events:
- Create a secondary SQS queue to act as the DLQ for your S3 event notification.
- In your S3 bucket's event notification configuration, set this DLQ as the "Dead-letter queue" under the SQS target settings.
- This way, any events that fail to reach your primary SQS queue will land in the DLQ, so you can manually process them later instead of losing them entirely.
Verify SQS Permissions for S3
Double-check your primary SQS queue's access policy to ensure S3 has permission to send messages to it. The policy should include a statement like this:
{ "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Principal": { "Service": "s3.amazonaws.com" }, "Action": "sqs:SendMessage", "Resource": "arn:aws:sqs:your-region:your-account-id:your-queue-name", "Condition": { "ArnLike": { "aws:SourceArn": "arn:aws:s3:::your-bucket-name" } } } ] }
If this policy is missing or misconfigured, S3 can't send events to SQS, especially under high load.
Tune S3 Event Notification Batching
By default, S3 batches events before sending them to SQS, but the default settings might not be optimal for high throughput. Adjust these values in your S3 event notification:
- Batch size: Increase from the default (100) to a higher value like 500 or 1000 (max is 1000). This reduces the number of API calls S3 makes to SQS, lowering the chance of throttling.
- Batch interval: Set to 5-10 seconds. This ensures S3 doesn't wait too long to send batches if the batch size isn't reached, while still reducing request volume.
2. Optimize SQS -> Lambda to Prevent Post-SQS Loss
Even once events reach SQS, we need to make sure Lambda processes them reliably, especially since you don't mind long processing times:
Configure a DLQ for the Lambda SQS Trigger
Set up a DLQ for your Lambda's SQS trigger. If Lambda fails to process a message (e.g., unhandled exception, timeout), the message will be retried a configurable number of times before being moved to the DLQ. This prevents messages from being lost in infinite retry loops or dropped silently:
- In your Lambda trigger settings, under "Dead-letter queue", select an SQS queue to use as the DLQ.
- Set the "Maximum receive count" (e.g., 3) — this is how many times Lambda will try to process a message before sending it to the DLQ.
Adjust Lambda Batch Size and Timeout
Your current batch size is 10, and each event takes ~35 seconds to process. Since your code processes events serially (using forEach), a batch of 10 will take ~350 seconds (almost 6 minutes) to complete. Make sure your Lambda's timeout is set to be longer than this — at least 10 minutes (600 seconds) to avoid timeout errors.
- If you want to reduce per-instance processing time, you could modify your Java code to process events in parallel (using
parallelStream()instead ofstream()), but that depends on whether your event processing is thread-safe. - Alternatively, lower the batch size to 1 if you want each Lambda instance to handle one event at a time — this will increase concurrency but ensure each event's processing time is predictable.
Check Lambda Concurrency Limits
If your Lambda function hits account-level concurrency limits, it won't be able to scale up to process all messages in SQS. Since you don't mind long processing times, you can either:
- Request a concurrency limit increase from AWS Support if needed.
- Set a reserved concurrency value for your function to ensure it gets enough capacity to handle the load.
Final Notes
With these changes, you should eliminate the 10% event loss: the S3 DLQ will catch any events that fail to reach the primary SQS queue, and the Lambda/SQS DLQ will ensure no messages are lost after reaching SQS. Your existing Lambda code looks fine for logging incoming events, so once the delivery pipeline is fixed, you should see all events in your logs.
内容的提问来源于stack exchange,提问作者user3549576

