如何使用pdf.js将PDF页面复制到目标PDF或新PDF并指定位置保存
Using pdf.js to Copy Pages Between PDF Documents
Hey there! Great question about copying pages between PDFs with pdf.js. First, a quick heads-up: pdf.js is designed mainly for parsing and rendering PDFs, not for creating or modifying them natively. To achieve page copying, we’ll combine pdf.js (for reading page data) with pdf-lib (a robust library for assembling and writing PDFs)—this is a common, reliable approach.
Prerequisites
First, include both libraries in your project. You can use CDNs for simplicity:
<script src="https://cdnjs.cloudflare.com/ajax/libs/pdf.js/4.0.379/pdf.min.js"></script> <script src="https://cdnjs.cloudflare.com/ajax/libs/pdf-lib/1.17.1/pdf-lib.min.js"></script>
Step-by-Step Implementation
Here’s a complete, reusable function that handles copying single or multiple pages, inserting them into an existing PDF or a new one, and saving the result:
async function copyPdfPages(sourcePdfUrl, targetPdfUrl = null, pagesToCopy, insertAt = -1) { try { // 1. Load the source PDF using pdf.js const sourcePdf = await pdfjsLib.getDocument(sourcePdfUrl).promise; // 2. Extract raw bytes for each page we want to copy const sourcePageBuffers = []; for (const pageNum of pagesToCopy) { // pdf.js uses 1-based page numbering const page = await sourcePdf.getPage(pageNum); // Get the raw page data as a Uint8Array const pageBytes = await sourcePdf.getPageRaw(pageNum); sourcePageBuffers.push(pageBytes); } // 3. Initialize the target document (existing or new) with pdf-lib const targetDoc = targetPdfUrl ? await PDFLib.PDFDocument.load(await fetch(targetPdfUrl).then(res => res.arrayBuffer())) : await PDFLib.PDFDocument.create(); // 4. Insert copied pages into the target document for (const pageBuffer of sourcePageBuffers) { // Load the single page as a temporary PDF document const tempDoc = await PDFLib.PDFDocument.load(pageBuffer); // Copy the page from the temp doc to the target const [copiedPage] = await targetDoc.copyPages(tempDoc, [0]); // Insert at specified position, or add to end if insertAt is -1 or out of bounds if (insertAt === -1 || insertAt >= targetDoc.getPageCount()) { targetDoc.addPage(copiedPage); } else { // pdf-lib uses 0-based indexing for insertion positions targetDoc.insertPage(insertAt, copiedPage); } } // 5. Save and download the modified PDF const finalPdfBytes = await targetDoc.save(); const blob = new Blob([finalPdfBytes], { type: 'application/pdf' }); const downloadUrl = URL.createObjectURL(blob); // Trigger download const link = document.createElement('a'); link.href = downloadUrl; link.download = 'modified-document.pdf'; link.click(); // Clean up URL.revokeObjectURL(downloadUrl); console.log("Pages copied successfully!"); } catch (error) { console.error("Error copying PDF pages:", error); } }
Example Usage
// Copy pages 1 and 3 from "source.pdf" to the end of "target.pdf" copyPdfPages("source.pdf", "target.pdf", [1, 3]); // Create a new PDF with just page 2 from "source.pdf" copyPdfPages("source.pdf", null, [2]); // Insert page 4 from "source.pdf" at position 2 (0-based) in "target.pdf" copyPdfPages("source.pdf", "target.pdf", [4], 2);
Key Notes
- Page Numbering: pdf.js uses 1-based page numbering (matches what users see), while pdf-lib uses 0-based for insertion positions.
- CORS: If loading PDFs from external domains, ensure the server allows cross-origin requests. Otherwise, host the PDFs on your own domain or use a proxy.
- Large Documents: For very large PDFs, process pages one at a time to avoid memory overload.
- Preserving Formatting: This method preserves most page content (text, images, layout) since we’re copying raw page data.
内容的提问来源于stack exchange,提问作者user1279887
相关产品推荐
相关产品推荐

