求助:Node.js Lambda提取S3中Raw格式邮件附件至另一S3桶
Hey there! Let’s walk through a solid plan to extract attachments from your raw SES emails stored in S3 and save them to another S3 bucket using your Node.js Lambda function. Since your existing pipeline is already working, we just need to extend the Lambda logic—here’s how to do it right:
First, update your Lambda’s IAM role to include the necessary S3 permissions:
- Allow
s3:GetObjecton your source S3 bucket (where raw emails are stored) - Allow
s3:PutObjecton your target S3 bucket (where attachments will go)
You can add these permissions via an inline policy in IAM, like this:
{ "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": "s3:GetObject", "Resource": "arn:aws:s3:::your-source-bucket/*" }, { "Effect": "Allow", "Action": "s3:PutObject", "Resource": "arn:aws:s3:::your-target-bucket/*" } ] }
Raw SES emails are in MIME format, so you’ll need a library to parse them and extract attachments. The most reliable choice for Node.js is mailparser—it handles all MIME edge cases and has a straightforward API.
You have two ways to include this dependency:
- Package with your code: Run
npm install mailparserlocally, then zip yourindex.jsandnode_modulesfolder for Lambda deployment. - Use a Lambda Layer: Create a layer containing the
mailparserpackage, then attach it to your Lambda. This keeps your deployment package small and reusable across functions.
Here’s a complete, tested Lambda function that handles the end-to-end process:
const AWS = require('aws-sdk'); const { simpleParser } = require('mailparser'); const s3 = new AWS.S3(); // Replace with your actual target bucket name const TARGET_BUCKET = 'your-target-attachments-bucket'; exports.handler = async (event) => { try { // Extract source bucket and object key from the S3 trigger event const s3Record = event.Records[0].s3; const sourceBucket = s3Record.bucket.name; // Decode URL-encoded object key (S3 encodes special characters like spaces) const sourceKey = decodeURIComponent(s3Record.object.key.replace(/\+/g, ' ')); // Fetch the raw email from S3 const emailResponse = await s3.getObject({ Bucket: sourceBucket, Key: sourceKey }).promise(); const rawEmailBuffer = emailResponse.Body; // Parse the raw email to extract attachments const parsedEmail = await simpleParser(rawEmailBuffer); // Upload each attachment to the target bucket for (const attachment of parsedEmail.attachments) { // Generate a unique key to avoid overwriting files // Customize this (e.g., use email message ID instead of timestamp) const attachmentKey = `extracted/${Date.now()}-${attachment.filename}`; await s3.putObject({ Bucket: TARGET_BUCKET, Key: attachmentKey, Body: attachment.content, ContentType: attachment.contentType, // Preserve original filename in Content-Disposition ContentDisposition: `attachment; filename="${attachment.filename}"` }).promise(); console.log(`Successfully saved attachment: ${attachmentKey}`); } return { statusCode: 200, body: JSON.stringify({ message: 'Attachments processed successfully' }) }; } catch (error) { console.error('Failed to process email attachments:', error); // Re-throw to trigger Lambda retries (or handle as needed) throw error; } };
- Lambda Configuration: Increase the function’s memory to at least 256MB (parsing large emails can be memory-intensive) and set the timeout to 10-30 seconds (to handle large uploads).
- Test with a Sample Email: Send a test email with attachments via SES, wait for it to land in your source bucket, then check if attachments appear in the target bucket.
- Debug with CloudWatch: If something breaks, check Lambda’s CloudWatch Logs for detailed errors (e.g., permission issues, parsing failures).
- Avoid Duplicate Uploads: Use the email’s
messageId(fromparsedEmail.messageId) instead of a timestamp in the attachment key—this ensures the same attachment isn’t uploaded multiple times if the Lambda retries. - Handle Large Attachments: For attachments larger than 512MB (Lambda’s memory limit), use S3 Multipart Upload. You can stream the attachment content directly from the parser to S3 to avoid loading the entire file into memory.
- Add a Dead-Letter Queue: Configure a SQS dead-letter queue for your Lambda to catch failed events, so you can retry processing without losing data.
内容的提问来源于stack exchange,提问作者Murillo Mamud

