使用Google Apps Script合并docx转PDF时解析失败求助
问题分析与解决方案
错误提示「No PDF header found」是因为你处理DOCX文件时,错误地将纯文本内容保存为PDF格式文件——这种文件只是后缀为PDF的纯文本,没有合法的PDF文件结构,导致pdf-lib无法解析。
核心问题点
- DOCX转PDF逻辑错误:你通过
DocumentApp提取文档纯文本,再手动创建PDF文件,这不是真正的PDF生成方式,生成的文件缺少PDF必需的头部信息。 - 无效ID未过滤:当URL解析失败时,
ids数组会包含空数组,后续data数组会出现undefined,可能引发额外错误。 sheet变量未定义:代码中直接使用sheet但未先获取表格对象。
修正后的完整代码
async function newMain() { // 1. 初始化必要对象(替换为你的表格名称和目标文件夹ID) const sheet = SpreadsheetApp.getActiveSpreadsheet().getSheetByName("Sheet1"); const destinationFolderId = "your-folder-id"; const destinationFolder = DriveApp.getFolderById(destinationFolderId); // 2. 解析文件URL并提取有效ID const urls = sheet.getRange(2, 2).getValue().toString().split(","); const validIds = urls.reduce((acc, url) => { const matches = url.match(/\/file\/d\/([^\/]+)\/edit/); if (matches && matches[1]) acc.push(matches[1]); return acc; }, []); // 3. 获取所有文件的PDF数据(DOCX自动转换为合法PDF) const pdfDataList = await Promise.all(validIds.map(async (id) => { const file = DriveApp.getFileById(id); const mimeType = file.getMimeType(); if (mimeType === 'application/vnd.openxmlformats-officedocument.wordprocessingml.document') { // DOCX转PDF:通过Drive API转换为Google文档,再导出为PDF const tempDocFile = Drive.Files.insert({}, file.getBlob(), { convert: true }); const tempDoc = DriveApp.getFileById(tempDocFile.id); const pdfBlob = tempDoc.getAs(MimeType.PDF); // 删除临时Google文档 Drive.Files.remove(tempDocFile.id); return new Uint8Array(pdfBlob.getBytes()); } else if (mimeType === 'application/pdf') { // 直接获取已有PDF的二进制数据 return new Uint8Array(file.getBlob().getBytes()); } else { // 跳过非PDF/DOCX文件 console.log(`跳过不支持的文件类型:${file.getName()}`); return null; } })); // 过滤掉无效的PDF数据 const validPdfData = pdfDataList.filter(data => data !== null); // 4. 加载pdf-lib并合并PDF const cdnjs = "https://cdn.jsdelivr.net/npm/pdf-lib/dist/pdf-lib.min.js"; eval(UrlFetchApp.fetch(cdnjs).getContentText().replace(/setTimeout\(.*?,.*?(\d*?)\)/g, "Utilities.sleep($1);return t();")); const pdfDoc = await PDFLib.PDFDocument.create(); for (const data of validPdfData) { const pdf = await PDFLib.PDFDocument.load(data); const pages = await pdfDoc.copyPages(pdf, pdf.getPageIndices()); pages.forEach(page => pdfDoc.addPage(page)); } // 5. 生成合并后的PDF文件 const mergedBytes = await pdfDoc.save(); const mergedFile = DriveApp.createFile( Utilities.newBlob([...new Int8Array(mergedBytes)], MimeType.PDF, "merged_result.pdf") ); mergedFile.moveTo(destinationFolder); }
关键修改说明
- DOCX转PDF逻辑修复:利用Drive API将DOCX转换为Google文档后,直接导出为合法的PDF Blob,确保生成的文件具备完整PDF结构。
- 无效ID过滤:使用
reduce收集有效的文件ID,避免后续处理出现undefined。 - 异步处理优化:用
Promise.all并行处理文件转换,提升效率;过滤无效数据,避免合并时出错。 - 变量初始化:补充
sheet变量的定义,确保代码可运行。
额外注意事项
- 确保已启用Drive API:在Google Apps Script编辑器中,依次点击「资源」→「高级Google服务」,找到「Drive API」并启用。
- 替换占位符:将代码中的
Sheet1(你的表格名称)和your-folder-id(目标文件夹ID)替换为实际值。 - 权限验证:首次运行时会要求授权,需允许脚本访问你的Google Drive和表格。
内容的提问来源于stack exchange,提问作者EagleEye
相关产品推荐
相关产品推荐

