You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何移除PDF中的所有图片?使用pdf-lib实现失败求助

用pdf-lib移除PDF所有图片的修复方案

你的代码没生效是因为这几个问题:

  • 只处理了第一页,未遍历PDF所有页面
  • 直接删除整个XObject资源字典,会误删非图片资源且破坏页面结构
  • 没有正确更新页面的资源配置逻辑

下面是能正常工作的代码:

const { PDFDocument, PDFName } = require('pdf-lib');

async function stripPdfImages(base64Pdf) {
  // 将Base64格式PDF转换为ArrayBuffer
  const byteArray = new Uint8Array(atob(base64Pdf).split('').map(char => char.charCodeAt(0)));
  const pdfDoc = await PDFDocument.load(byteArray.buffer);

  // 遍历所有页面处理图片资源
  for (const page of pdfDoc.getPages()) {
    const resources = page.node.Resources();
    if (!resources) continue;

    // 获取页面的XObject资源(不存在则跳过)
    const xObjects = resources.lookupMaybe(PDFName.of('XObject'));
    if (!xObjects?.entries) continue;

    // 过滤保留非图片类型的XObject
    const keepXObjects = new Map();
    for (const [key, obj] of xObjects.entries()) {
      const subtype = obj.lookup(PDFName.of('Subtype'));
      if (subtype !== PDFName.of('Image')) {
        keepXObjects.set(key, obj);
      }
    }

    // 更新页面的XObject资源字典
    resources.set(PDFName.of('XObject'), pdfDoc.context.obj(keepXObjects));
  }

  // 导出处理后的PDF字节流
  return await pdfDoc.save();
}

// 调用示例
// const cleanedPdfBytes = await stripPdfImages(window.pdfFile);

关键修复说明:

  • 遍历PDF所有页面,确保每页的图片都被处理
  • 精准过滤:仅剔除Subtype为Image的资源,保留表单、矢量图等其他合法XObject
  • 安全更新:用过滤后的资源字典替换原字典,避免直接删除导致的页面渲染异常

内容的提问来源于stack exchange,提问作者mobiw

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 17:17:09