使用pdf.js渲染服务器PDF页面时为何出现内存泄漏?
Hey there! Let’s dig into why your custom implementation is hitting 500MB of memory while the pre-built version stays at 100-200MB when rendering those 180 individual PDF pages. Here are the most likely culprits:
1. Uncleaned PDF Document Instances
Every time you call PDFJS.getDocument({ data: atob(pageData) }), you’re creating a new PDFDocumentProxy instance. If you don’t explicitly destroy these instances after rendering, the browser’s garbage collector (GC) might not clean them up—especially for pages with complex math formulas and charts, which hold onto large resources like fonts, rasterized graphics, and parsed data.
The pre-built version almost certainly handles this automatically. Fix it by calling destroy() on the document once you’re done with the page:
PDFJS.getDocument({ data: atob(pageData) }).then(pdfDoc => { return pdfDoc.getPage(1).then(page => { // Your page rendering logic here const canvas = document.getElementById('your-canvas'); const viewport = page.getViewport({ scale: 1.0 }); canvas.height = viewport.height; canvas.width = viewport.width; return page.render({ canvasContext: canvas.getContext('2d'), viewport: viewport }).promise.then(() => { // Destroy the document to free memory pdfDoc.destroy(); }); }); });
2. Lingering DOM/Canvas References
If you’re rendering each page to a canvas or DOM element that’s not properly removed from the DOM (or still has active references—like global variables or unremoved event listeners), those elements will hang around in memory. Complex charts and formula-heavy pages have larger canvas buffers, so 180 of them add up fast.
Pre-built versions typically have a cleanup flow that removes old canvases from the DOM and nulls out references to let GC do its job. Make sure you’re:
- Removing unused canvases from the DOM after rendering the next page
- Nulling out any variables that reference the canvas or page objects
3. Inefficient Data Handling with atob()
Calling atob(pageData) creates a new string object for each page. For large pages (especially those with images/charts), these strings can be massive, and if they’re not immediately garbage collected, they pile up. The pre-built version probably uses more efficient binary data formats like ArrayBuffer instead of raw strings, which reduces memory overhead and lets pdf.js process data faster.
Try converting your base64 data to a Uint8Array before passing it to pdf.js:
// Convert base64 to Uint8Array const binaryString = atob(pageData); const uint8Array = new Uint8Array(binaryString.length); for (let i = 0; i < binaryString.length; i++) { uint8Array[i] = binaryString.charCodeAt(i); } // Use the Uint8Array instead of the decoded string PDFJS.getDocument({ data: uint8Array }).then(...);
4. No Throttling for GC
If you’re loading all 180 pages in a tight loop, the browser’s GC doesn’t get a chance to run between renders. The pre-built version likely uses batch loading or throttling (e.g., loading 20 pages, waiting a few milliseconds, then loading the next batch) to give GC time to clean up old resources before creating new ones.
Try adding a small delay between page loads to let GC catch up:
async function loadPages(pageDataList) { for (const pageData of pageDataList) { await renderSinglePage(pageData); // Your existing render function // Add a small delay to allow GC to run await new Promise(resolve => setTimeout(resolve, 50)); } }
5. Unoptimized pdf.js Configuration
Pre-built pdf.js builds often come with optimizations disabled for features you don’t need (like text extraction, annotations, or font subsetting). If your custom code uses the default configuration, it might be loading extra functionality that eats up memory.
Try tweaking these config options to reduce overhead:
PDFJS.getDocument({ data: ..., disableFontFace: true, // Disable if you don't need selectable text disableStream: true, disableAutoFetch: true, maxCanvasPixels: 4096 * 4096 // Limit canvas size if possible }).then(...);
Start with checking if you’re destroying document instances and cleaning up DOM references—those are the most common fixes for memory bloat in pdf.js implementations.
内容的提问来源于stack exchange,提问作者sofvlad

