在AWS Lambda中用Python实现HTML字符串转PDF并转Base64发送
AWS Lambda中实现HTML转PDF并转Base64的可行性与实现方案
完全可行。AWS Lambda提供了**/tmp临时目录**用于文件读写,当前该目录最大容量为10GB,足够存放临时HTML和PDF文件,完全能支撑你描述的流程。
具体实现步骤
- 请求外部源1,获取包含HTML字符串的响应数据
- 将HTML字符串写入Lambda的
/tmp/Sample.html文件(Lambda仅/tmp目录具备可写权限,其他系统目录均为只读) - 使用HTML转PDF工具(如Puppeteer、wkhtmltopdf)将/tmp下的HTML文件转换为PDF,输出到
/tmp/Sample.pdf - 读取PDF文件内容,转换为Base64编码字符串
- 将Base64字符串发送至外部源2
Node.js代码示例
以下是基于Puppeteer的实现示例(需配合chrome-aws-lambda层使用,避免打包过大的Chrome二进制文件):
const axios = require('axios'); const puppeteer = require('puppeteer-core'); const chrome = require('chrome-aws-lambda'); const fs = require('fs').promises; const path = require('path'); exports.handler = async (event) => { try { // 1. 请求外部源1获取HTML字符串 const externalRes = await axios.get('https://外部源1的接口地址'); const htmlContent = externalRes.data; // 2. 将HTML写入/tmp目录 const htmlPath = path.join('/tmp', 'Sample.html'); await fs.writeFile(htmlPath, htmlContent, 'utf8'); // 3. 用Puppeteer将HTML转PDF const browser = await puppeteer.launch({ args: chrome.args, executablePath: await chrome.executablePath, headless: chrome.headless, }); const page = await browser.newPage(); await page.goto(`file://${htmlPath}`, { waitUntil: 'networkidle0' }); const pdfPath = path.join('/tmp', 'Sample.pdf'); await page.pdf({ path: pdfPath, format: 'A4' }); await browser.close(); // 4. 将PDF转换为Base64 const pdfBuffer = await fs.readFile(pdfPath); const pdfBase64 = pdfBuffer.toString('base64'); // 5. 发送至外部源2 await axios.post('https://外部源2的接口地址', { pdf: pdfBase64 }); return { statusCode: 200, body: '处理完成' }; } catch (error) { console.error('处理失败:', error); return { statusCode: 500, body: '处理失败' }; } };
注意事项
- 必须使用
/tmp目录存储临时文件,Lambda的其他目录均为只读 - 依赖管理:若使用Puppeteer,建议通过Lambda层引入
chrome-aws-lambda和puppeteer-core,避免将Chrome二进制文件打包进部署包导致体积过大 - 资源配置:转PDF操作需要一定内存,建议将Lambda的内存配置至少设为512MB,同时根据实际情况调整超时时间(最长可设为15分钟)
- 清理临时文件:虽然Lambda执行结束后/tmp目录会被清理,但如果同一容器被复用,手动删除临时文件可避免占用空间
内容的提问来源于stack exchange,提问作者That Coder
相关产品推荐
相关产品推荐

