You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Node.js中使用PDFkit生成PDF时,如何解决HTML标签渲染失效问题?

基于PDFkit渲染HTML内容的可行方案

针对你用PDFkit生成PDF时HTML格式丢失的问题,以下是三个可行方案,按实用性和复杂度排序:

方案2:使用html-to-pdfkit库(推荐)

这个库专为PDFkit设计,能直接将HTML转换为PDFkit的绘制指令,支持大部分常用HTML标签(strong、em、h1-h6、p、ul/ol等)和基础样式,无需手动解析标签,集成成本极低。

步骤

  1. 安装依赖:
npm install html-to-pdfkit
  1. 示例代码:
const PDFDocument = require('pdfkit');
const htmlToPDFKit = require('html-to-pdfkit');
const fs = require('fs');

const doc = new PDFDocument();
const htmlData = `<!DOCTYPE html>
<html>
<body>
<h1>The h1 element</h1>
<p>This is normal text - <em><strong>and this is bold italic text</strong></em>.</p>
<ol>
<li>Coffee</li>
<li>Tea</li>
<li>Milk</li>
</ol>
<ul>
<li>Coffee</li>
<li>Tea</li>
<li>Milk</li>
</ul>
</body>
</html>`;

// 将HTML渲染到PDF文档,指定宽度和对齐方式
htmlToPDFKit(doc, htmlData, {
  width: 500,
  align: 'left'
});

doc.pipe(fs.createWriteStream('./html-rendered.pdf'));
doc.end();

优缺点

  • 优点:无缝对接PDFkit,开发成本低;支持基础HTML格式和样式;保持PDF文档连贯性。
  • 缺点:复杂CSS(如浮动、自定义布局)支持有限;依赖第三方库,需关注版本兼容。

方案1:手动解析HTML标签(基于Cheerio)

用Cheerio解析HTML结构,遍历节点并调用PDFkit对应API设置样式,适合需要高度定制PDF样式、HTML标签类型有限的场景。

步骤

  1. 安装依赖:
npm install cheerio
  1. 示例代码:
const PDFDocument = require('pdfkit');
const cheerio = require('cheerio');
const fs = require('fs');

const doc = new PDFDocument();
const htmlData = `<!DOCTYPE html>
<html>
<body>
<h1>The h1 element</h1>
<p>This is normal text - <em><strong>and this is bold italic text</strong></em>.</p>
<ol>
<li>Coffee</li>
<li>Tea</li>
<li>Milk</li>
</ol>
<ul>
<li>Coffee</li>
<li>Tea</li>
<li>Milk</li>
</ul>
</body>
</html>`;

const $ = cheerio.load(htmlData);

// 遍历body下的所有元素
$('body').children().each((_, element) => {
  const tag = element.tagName.toLowerCase();
  
  switch(tag) {
    case 'h1':
      doc.fontSize(24).font('Helvetica-Bold').text($(element).text().trim());
      doc.fontSize(12).font('Helvetica').moveDown(0.5); // 恢复默认样式
      break;
    case 'p':
      processInlineElements($(element), doc);
      doc.moveDown(0.5);
      break;
    case 'ol':
      $(element).find('li').each((index, li) => {
        doc.text(`${index + 1}. ${$(li).text().trim()}`, { indent: 20 });
      });
      doc.moveDown(0.5);
      break;
    case 'ul':
      $(element).find('li').each((_, li) => {
        doc.text(`• ${$(li).text().trim()}`, { indent: 20 });
      });
      doc.moveDown(0.5);
      break;
  }
});

// 处理行内样式(strong、em)
function processInlineElements($element, doc) {
  let currentFont = 'Helvetica';
  let isItalic = false;
  
  $element.contents().each((_, node) => {
    if (node.type === 'text') {
      doc.font(currentFont).fontItalic(isItalic).text(node.data.trim(), { continued: true });
    } else if (node.type === 'tag') {
      const tag = node.tagName.toLowerCase();
      if (tag === 'strong') {
        currentFont = 'Helvetica-Bold';
        processInlineElements($(node), doc);
        currentFont = 'Helvetica';
      } else if (tag === 'em') {
        isItalic = true;
        processInlineElements($(node), doc);
        isItalic = false;
      }
    }
  });
}

doc.pipe(fs.createWriteStream('./manual-parse.pdf'));
doc.end();

优缺点

  • 优点:完全基于PDFkit,样式定制精度高;无重型依赖。
  • 缺点:需手动处理所有支持的标签,复杂HTML开发成本高。

方案3:复杂HTML转PDF后合并(基于Puppeteer + pdf-lib)

如果HTML包含复杂CSS、动态内容或复杂布局,可先用Puppeteer将HTML转成独立PDF,再用pdf-lib合并到PDFkit生成的主文档中。

步骤

  1. 安装依赖:
npm install puppeteer pdf-lib
  1. 示例代码:
const PDFDocument = require('pdfkit');
const puppeteer = require('puppeteer');
const { PDFDocument: PDFFromLib } = require('pdf-lib');
const fs = require('fs').promises;
const fsSync = require('fs');

async function generateMergedPDF() {
  // 1. 用PDFkit生成主文档
  const mainDoc = new PDFDocument();
  const mainStream = fsSync.createWriteStream('./main-temp.pdf');
  mainDoc.text('PDFkit生成的主文档内容', { align: 'left' });
  mainDoc.moveDown(1);
  mainDoc.end();
  await new Promise(resolve => mainStream.on('finish', resolve));

  // 2. 用Puppeteer将HTML转成PDF片段
  const browser = await puppeteer.launch();
  const page = await browser.newPage();
  const htmlData = `<!DOCTYPE html>
<html>
<body>
<h1>The h1 element</h1>
<p>This is normal text - <em><strong>and this is bold italic text</strong></em>.</p>
<ol>
<li>Coffee</li>
<li>Tea</li>
<li>Milk</li>
</ol>
<ul>
<li>Coffee</li>
<li>Tea</li>
<li>Milk</li>
</ul>
</body>
</html>`;
  await page.setContent(htmlData);
  const htmlPdfBytes = await page.pdf({ format: 'A4' });
  await browser.close();

  // 3. 合并两个PDF
  const mainPdfBytes = await fs.readFile('./main-temp.pdf');
  const mergedPdf = await PDFFromLib.create();
  
  const mainPdf = await PDFFromLib.load(mainPdfBytes);
  const htmlPdf = await PDFFromLib.load(htmlPdfBytes);
  
  // 复制主文档页面
  const mainPages = await mergedPdf.copyPages(mainPdf, mainPdf.getPageIndices());
  mainPages.forEach(page => mergedPdf.addPage(page));
  
  // 复制HTML生成的页面
  const htmlPages = await mergedPdf.copyPages(htmlPdf, htmlPdf.getPageIndices());
  htmlPages.forEach(page => mergedPdf.addPage(page));
  
  // 保存合并后的PDF
  const mergedBytes = await mergedPdf.save();
  await fs.writeFile('./merged.pdf', mergedBytes);
  
  // 删除临时文件
  await fs.unlink('./main-temp.pdf');
}

generateMergedPDF();

优缺点

  • 优点:支持所有浏览器可渲染的HTML/CSS,包括复杂布局和动态内容。
  • 缺点:依赖Puppeteer(需Chromium),体积大、启动慢;合并流程复杂,性能不如前两个方案。

最佳选择建议

  • 基础HTML格式(标题、段落、粗体/斜体、列表):优先选方案2,开发效率最高。
  • 需要高度定制PDF样式:选方案1,可控性强。
  • 复杂HTML/CSS布局:选方案3,兼容性最好。

内容的提问来源于stack exchange,提问作者Rajasekhar Reddy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 10:20:17