You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Google Apps Script合并docx转PDF时解析失败求助

问题分析与解决方案

错误提示「No PDF header found」是因为你处理DOCX文件时,错误地将纯文本内容保存为PDF格式文件——这种文件只是后缀为PDF的纯文本,没有合法的PDF文件结构,导致pdf-lib无法解析。

核心问题点

  1. DOCX转PDF逻辑错误:你通过DocumentApp提取文档纯文本,再手动创建PDF文件,这不是真正的PDF生成方式,生成的文件缺少PDF必需的头部信息。
  2. 无效ID未过滤:当URL解析失败时,ids数组会包含空数组,后续data数组会出现undefined,可能引发额外错误。
  3. sheet变量未定义:代码中直接使用sheet但未先获取表格对象。

修正后的完整代码

async function newMain() {
  // 1. 初始化必要对象(替换为你的表格名称和目标文件夹ID)
  const sheet = SpreadsheetApp.getActiveSpreadsheet().getSheetByName("Sheet1");
  const destinationFolderId = "your-folder-id";
  const destinationFolder = DriveApp.getFolderById(destinationFolderId);

  // 2. 解析文件URL并提取有效ID
  const urls = sheet.getRange(2, 2).getValue().toString().split(",");
  const validIds = urls.reduce((acc, url) => {
    const matches = url.match(/\/file\/d\/([^\/]+)\/edit/);
    if (matches && matches[1]) acc.push(matches[1]);
    return acc;
  }, []);

  // 3. 获取所有文件的PDF数据(DOCX自动转换为合法PDF)
  const pdfDataList = await Promise.all(validIds.map(async (id) => {
    const file = DriveApp.getFileById(id);
    const mimeType = file.getMimeType();

    if (mimeType === 'application/vnd.openxmlformats-officedocument.wordprocessingml.document') {
      // DOCX转PDF:通过Drive API转换为Google文档,再导出为PDF
      const tempDocFile = Drive.Files.insert({}, file.getBlob(), { convert: true });
      const tempDoc = DriveApp.getFileById(tempDocFile.id);
      const pdfBlob = tempDoc.getAs(MimeType.PDF);
      
      // 删除临时Google文档
      Drive.Files.remove(tempDocFile.id);
      
      return new Uint8Array(pdfBlob.getBytes());
    } else if (mimeType === 'application/pdf') {
      // 直接获取已有PDF的二进制数据
      return new Uint8Array(file.getBlob().getBytes());
    } else {
      // 跳过非PDF/DOCX文件
      console.log(`跳过不支持的文件类型:${file.getName()}`);
      return null;
    }
  }));

  // 过滤掉无效的PDF数据
  const validPdfData = pdfDataList.filter(data => data !== null);

  // 4. 加载pdf-lib并合并PDF
  const cdnjs = "https://cdn.jsdelivr.net/npm/pdf-lib/dist/pdf-lib.min.js";
  eval(UrlFetchApp.fetch(cdnjs).getContentText().replace(/setTimeout\(.*?,.*?(\d*?)\)/g, "Utilities.sleep($1);return t();"));

  const pdfDoc = await PDFLib.PDFDocument.create();
  for (const data of validPdfData) {
    const pdf = await PDFLib.PDFDocument.load(data);
    const pages = await pdfDoc.copyPages(pdf, pdf.getPageIndices());
    pages.forEach(page => pdfDoc.addPage(page));
  }

  // 5. 生成合并后的PDF文件
  const mergedBytes = await pdfDoc.save();
  const mergedFile = DriveApp.createFile(
    Utilities.newBlob([...new Int8Array(mergedBytes)], MimeType.PDF, "merged_result.pdf")
  );
  mergedFile.moveTo(destinationFolder);
}

关键修改说明

  • DOCX转PDF逻辑修复:利用Drive API将DOCX转换为Google文档后,直接导出为合法的PDF Blob,确保生成的文件具备完整PDF结构。
  • 无效ID过滤:使用reduce收集有效的文件ID,避免后续处理出现undefined。
  • 异步处理优化:用Promise.all并行处理文件转换,提升效率;过滤无效数据,避免合并时出错。
  • 变量初始化:补充sheet变量的定义,确保代码可运行。

额外注意事项

  1. 确保已启用Drive API:在Google Apps Script编辑器中,依次点击「资源」→「高级Google服务」,找到「Drive API」并启用。
  2. 替换占位符:将代码中的Sheet1(你的表格名称)和your-folder-id(目标文件夹ID)替换为实际值。
  3. 权限验证:首次运行时会要求授权,需允许脚本访问你的Google Drive和表格。

内容的提问来源于stack exchange,提问作者EagleEye

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 11:35:06