如何用pdfjs库获取PDF阅读器填写的数据?跨域问题求助
解决iframe中PDF阅读器填写数据获取的跨域问题
问题原因分析
- 你遇到的跨域错误本质是:浏览器内置PDF阅读器的
contentWindow属于不同源上下文,无法直接调用其fetch方法。 - 更核心的问题:即便没有跨域限制,
iframe.src的data URI也不会同步用户填写的表单数据——浏览器内置阅读器仅在内存中渲染和临时存储表单内容,不会修改原始的PDF源文件。
解决方案
放弃使用浏览器内置PDF阅读器,改用PDF.js接管PDF的渲染与表单处理,通过其API直接访问用户填写的表单数据,或导出包含填写内容的完整PDF文件。
完整代码示例
<html> <head> <meta charset="utf-8" /> <script src="https://unpkg.com/pdf-lib@1.17.1/dist/pdf-lib.min.js"></script> <script src="https://unpkg.com/downloadjs@1.4.7"></script> <!-- 引入PDF.js核心库 --> <script src="https://cdnjs.cloudflare.com/ajax/libs/pdf.js/3.11.174/pdf.min.js"></script> <style> #pdf-container { width: 100%; height: 80vh; border: 1px solid #ccc; margin-top: 10px; } </style> </head> <body> <button onclick="fillForm()">Fill PDF</button> <button onclick="getNewPDF()">Get Updated PDF</button> <div id="pdf-container"></div> <script> const { PDFDocument } = PDFLib; let pdfViewer = null; let pdfDoc = null; // 初始化PDF.js渲染器,加载并展示PDF async function initPdfViewer(pdfBytes) { const container = document.getElementById('pdf-container'); const pdfjsLib = window['pdfjs-dist/build/pdf']; // 配置PDF.js的worker脚本地址 pdfjsLib.GlobalWorkerOptions.workerSrc = 'https://cdnjs.cloudflare.com/ajax/libs/pdf.js/3.11.174/pdf.worker.min.js'; // 加载PDF文档 pdfDoc = await pdfjsLib.getDocument({ data: pdfBytes }).promise; // 创建PDF查看器实例 const viewer = new pdfjsLib.PDFViewer({ container: container, }); viewer.setDocument(pdfDoc); pdfViewer = viewer; // 等待文档完全加载 await pdfDoc.promise; } async function fillForm() { // 拉取原始PDF文件 const formUrl = 'http://localhost:3000/pdfFile.pdf'; const formPdfBytes = await fetch(formUrl, { method: 'get', mode: 'cors', headers: { 'Content-Type': 'application/pdf' } }).then(res => res.arrayBuffer()); // 用pdf-lib填充初始表单字段 const pdfDocLib = await PDFDocument.load(formPdfBytes); const form = pdfDocLib.getForm(); const nameField = form.getTextField('name'); nameField.setText('Shoaib Raza'); const addressField = form.getTextField('address'); addressField.setText('Hello World'); // 导出填充后的PDF字节数据 const updatedPdfBytes = await pdfDocLib.save(); // 交给PDF.js渲染展示 await initPdfViewer(updatedPdfBytes); } async function getNewPDF(){ if (!pdfDoc) return; // 方式1:获取所有表单字段的当前填写值 const formFields = await pdfDoc.getFieldObjects(); const fieldValues = {}; for (const [fieldName, field] of Object.entries(formFields)) { fieldValues[fieldName] = await field.getValue(); } console.log('表单字段填写值:', fieldValues); // 方式2:导出包含填写内容的完整PDF文件 const originalPdfUrl = 'http://localhost:3000/pdfFile.pdf'; const originalPdfBytes = await fetch(originalPdfUrl).then(res => res.arrayBuffer()); const pdfDocLib = await PDFDocument.load(originalPdfBytes); const form = pdfDocLib.getForm(); // 用用户填写的字段值更新PDF for (const [fieldName, value] of Object.entries(fieldValues)) { const field = form.getField(fieldName); if (field instanceof PDFLib.TextField) { field.setText(value); } // 可扩展处理复选框、单选框等其他类型字段 } // 导出最终PDF const finalPdfBytes = await pdfDocLib.save(); const finalPdfBase64 = await pdfDocLib.saveAsBase64(); console.log('最终PDF Base64:', finalPdfBase64); // 可选:触发下载 download(finalPdfBytes, 'updated-form.pdf', 'application/pdf'); } </script> </body> </html>
关键说明
- PDF.js的优势:作为浏览器端原生运行的PDF处理库,它完全绕过了浏览器内置阅读器的同源限制,能直接访问并操作PDF表单数据。
- 两种数据获取逻辑:
- 仅获取字段键值对:适合只需收集表单数据的场景,无需生成完整PDF。
- 生成完整PDF:结合pdf-lib重新渲染PDF,确保导出文件包含所有填写内容。
- 注意事项:
- 确保PDF.js主库与worker脚本版本一致,避免兼容性问题。
- 若PDF包含复选框、单选框等复杂字段,需扩展
getNewPDF中的字段处理逻辑。 - 测试时保证
http://localhost:3000的CORS配置允许当前页面访问。
内容的提问来源于stack exchange,提问作者Shoaib Raza
相关产品推荐
相关产品推荐

