如何用Node Lambda提取S3中SES邮件的text/plain内容?方案选择
Great question! When dealing with email content pulled from SES/S3, you’ve got two solid options—let’s break them down, along with their pros and cons.
Option 1: Use an NPM Email Parsing Library (Recommended)
Regex works for simple cases, but email formats are surprisingly messy (think multipart messages, encoded content, varying boundary markers, etc.). Using a dedicated library like mailparser will save you from headaches down the line.
Here’s how to implement this:
- Install the dependency locally:
npm install mailparser - Package your Lambda code + dependencies into a ZIP (or use a Lambda Layer for reusability across multiple functions).
- Update your Lambda function:
const { simpleParser } = require('mailparser'); const AWS = require('aws-sdk'); const s3 = new AWS.S3(); const bucketName = 'your-bucket-name-here'; // Don't forget to set your actual bucket name! exports.handler = async function(event, context) { try { const sesNotification = event.Records[0].ses; const s3Response = await s3.getObject({ Bucket: bucketName, Key: `ses/${sesNotification.mail.messageId}` }).promise(); // Parse the raw email content const parsedEmail = await simpleParser(s3Response.Body); const plainTextContent = parsedEmail.text; // This directly gives you the text/plain section console.log('Extracted plain text:', plainTextContent); // Add your logic to use the plainTextContent here return { statusCode: 200, body: 'Content extracted successfully' }; } catch (err) { console.error('Error during processing:', err); throw err; } };
This approach automatically handles edge cases like base64 encoding, multipart boundaries, and nested content—way more reliable than regex for real-world emails.
Option 2: Use Regular Expressions (For Simple Scenarios)
If you’re working with extremely straightforward emails and want to avoid adding dependencies, regex can work—but be warned, it’s fragile. Here’s a basic pattern to target the text/plain section:
const AWS = require('aws-sdk'); const s3 = new AWS.S3(); const bucketName = 'your-bucket-name-here'; exports.handler = async function(event, context) { try { const sesNotification = event.Records[0].ses; const s3Response = await s3.getObject({ Bucket: bucketName, Key: `ses/${sesNotification.mail.messageId}` }).promise(); const emailContent = s3Response.Body.toString(); // Regex to match text/plain content (handles common boundary cases) const plainTextRegex = /Content-Type: text\/plain[\s\S]+?\r\n\r\n([\s\S]+?)(?=\r\n--|$)/i; const match = emailContent.match(plainTextRegex); let plainTextContent = null; if (match && match[1]) { // Decode base64 if the email uses that encoding if (emailContent.includes('Content-Transfer-Encoding: base64')) { plainTextContent = Buffer.from(match[1].trim(), 'base64').toString('utf8'); } else { plainTextContent = match[1].trim(); } } console.log('Extracted plain text:', plainTextContent); // Proceed with your custom logic here return { statusCode: 200, body: 'Content extracted successfully' }; } catch (err) { console.error('Error during processing:', err); throw err; } };
Caveats with Regex:
- Won’t handle nested multipart messages well.
- Fails if the email uses non-standard boundary markers or encoding types like quoted-printable.
- Requires constant updates if you encounter new email formatting variations.
Final Recommendation
Stick with the NPM library approach unless you have a very specific reason to avoid dependencies. It’s more maintainable and far less likely to break when faced with unexpected email formatting.
内容的提问来源于stack exchange,提问作者Ben Hall

