Node.js中使用PDFkit生成PDF时,如何解决HTML标签渲染失效问题?
基于PDFkit渲染HTML内容的可行方案
针对你用PDFkit生成PDF时HTML格式丢失的问题,以下是三个可行方案,按实用性和复杂度排序:
方案2:使用html-to-pdfkit库(推荐)
这个库专为PDFkit设计,能直接将HTML转换为PDFkit的绘制指令,支持大部分常用HTML标签(strong、em、h1-h6、p、ul/ol等)和基础样式,无需手动解析标签,集成成本极低。
步骤
- 安装依赖:
npm install html-to-pdfkit
- 示例代码:
const PDFDocument = require('pdfkit'); const htmlToPDFKit = require('html-to-pdfkit'); const fs = require('fs'); const doc = new PDFDocument(); const htmlData = `<!DOCTYPE html> <html> <body> <h1>The h1 element</h1> <p>This is normal text - <em><strong>and this is bold italic text</strong></em>.</p> <ol> <li>Coffee</li> <li>Tea</li> <li>Milk</li> </ol> <ul> <li>Coffee</li> <li>Tea</li> <li>Milk</li> </ul> </body> </html>`; // 将HTML渲染到PDF文档,指定宽度和对齐方式 htmlToPDFKit(doc, htmlData, { width: 500, align: 'left' }); doc.pipe(fs.createWriteStream('./html-rendered.pdf')); doc.end();
优缺点
- 优点:无缝对接PDFkit,开发成本低;支持基础HTML格式和样式;保持PDF文档连贯性。
- 缺点:复杂CSS(如浮动、自定义布局)支持有限;依赖第三方库,需关注版本兼容。
方案1:手动解析HTML标签(基于Cheerio)
用Cheerio解析HTML结构,遍历节点并调用PDFkit对应API设置样式,适合需要高度定制PDF样式、HTML标签类型有限的场景。
步骤
- 安装依赖:
npm install cheerio
- 示例代码:
const PDFDocument = require('pdfkit'); const cheerio = require('cheerio'); const fs = require('fs'); const doc = new PDFDocument(); const htmlData = `<!DOCTYPE html> <html> <body> <h1>The h1 element</h1> <p>This is normal text - <em><strong>and this is bold italic text</strong></em>.</p> <ol> <li>Coffee</li> <li>Tea</li> <li>Milk</li> </ol> <ul> <li>Coffee</li> <li>Tea</li> <li>Milk</li> </ul> </body> </html>`; const $ = cheerio.load(htmlData); // 遍历body下的所有元素 $('body').children().each((_, element) => { const tag = element.tagName.toLowerCase(); switch(tag) { case 'h1': doc.fontSize(24).font('Helvetica-Bold').text($(element).text().trim()); doc.fontSize(12).font('Helvetica').moveDown(0.5); // 恢复默认样式 break; case 'p': processInlineElements($(element), doc); doc.moveDown(0.5); break; case 'ol': $(element).find('li').each((index, li) => { doc.text(`${index + 1}. ${$(li).text().trim()}`, { indent: 20 }); }); doc.moveDown(0.5); break; case 'ul': $(element).find('li').each((_, li) => { doc.text(`• ${$(li).text().trim()}`, { indent: 20 }); }); doc.moveDown(0.5); break; } }); // 处理行内样式(strong、em) function processInlineElements($element, doc) { let currentFont = 'Helvetica'; let isItalic = false; $element.contents().each((_, node) => { if (node.type === 'text') { doc.font(currentFont).fontItalic(isItalic).text(node.data.trim(), { continued: true }); } else if (node.type === 'tag') { const tag = node.tagName.toLowerCase(); if (tag === 'strong') { currentFont = 'Helvetica-Bold'; processInlineElements($(node), doc); currentFont = 'Helvetica'; } else if (tag === 'em') { isItalic = true; processInlineElements($(node), doc); isItalic = false; } } }); } doc.pipe(fs.createWriteStream('./manual-parse.pdf')); doc.end();
优缺点
- 优点:完全基于PDFkit,样式定制精度高;无重型依赖。
- 缺点:需手动处理所有支持的标签,复杂HTML开发成本高。
方案3:复杂HTML转PDF后合并(基于Puppeteer + pdf-lib)
如果HTML包含复杂CSS、动态内容或复杂布局,可先用Puppeteer将HTML转成独立PDF,再用pdf-lib合并到PDFkit生成的主文档中。
步骤
- 安装依赖:
npm install puppeteer pdf-lib
- 示例代码:
const PDFDocument = require('pdfkit'); const puppeteer = require('puppeteer'); const { PDFDocument: PDFFromLib } = require('pdf-lib'); const fs = require('fs').promises; const fsSync = require('fs'); async function generateMergedPDF() { // 1. 用PDFkit生成主文档 const mainDoc = new PDFDocument(); const mainStream = fsSync.createWriteStream('./main-temp.pdf'); mainDoc.text('PDFkit生成的主文档内容', { align: 'left' }); mainDoc.moveDown(1); mainDoc.end(); await new Promise(resolve => mainStream.on('finish', resolve)); // 2. 用Puppeteer将HTML转成PDF片段 const browser = await puppeteer.launch(); const page = await browser.newPage(); const htmlData = `<!DOCTYPE html> <html> <body> <h1>The h1 element</h1> <p>This is normal text - <em><strong>and this is bold italic text</strong></em>.</p> <ol> <li>Coffee</li> <li>Tea</li> <li>Milk</li> </ol> <ul> <li>Coffee</li> <li>Tea</li> <li>Milk</li> </ul> </body> </html>`; await page.setContent(htmlData); const htmlPdfBytes = await page.pdf({ format: 'A4' }); await browser.close(); // 3. 合并两个PDF const mainPdfBytes = await fs.readFile('./main-temp.pdf'); const mergedPdf = await PDFFromLib.create(); const mainPdf = await PDFFromLib.load(mainPdfBytes); const htmlPdf = await PDFFromLib.load(htmlPdfBytes); // 复制主文档页面 const mainPages = await mergedPdf.copyPages(mainPdf, mainPdf.getPageIndices()); mainPages.forEach(page => mergedPdf.addPage(page)); // 复制HTML生成的页面 const htmlPages = await mergedPdf.copyPages(htmlPdf, htmlPdf.getPageIndices()); htmlPages.forEach(page => mergedPdf.addPage(page)); // 保存合并后的PDF const mergedBytes = await mergedPdf.save(); await fs.writeFile('./merged.pdf', mergedBytes); // 删除临时文件 await fs.unlink('./main-temp.pdf'); } generateMergedPDF();
优缺点
- 优点:支持所有浏览器可渲染的HTML/CSS,包括复杂布局和动态内容。
- 缺点:依赖Puppeteer(需Chromium),体积大、启动慢;合并流程复杂,性能不如前两个方案。
最佳选择建议
- 基础HTML格式(标题、段落、粗体/斜体、列表):优先选方案2,开发效率最高。
- 需要高度定制PDF样式:选方案1,可控性强。
- 复杂HTML/CSS布局:选方案3,兼容性最好。
内容的提问来源于stack exchange,提问作者Rajasekhar Reddy
相关产品推荐
相关产品推荐

