Ionic3+Angular5应用使用pdfreader解析PDF时出现fs.readFileSync is not a function错误的解决方案咨询
Hey there, let's break down why you're seeing that error and get your PDF text extraction working smoothly!
Why the Error Happens
The pdfreader package you’re using is built for Node.js environments, where it relies on Node’s fs (file system) module to read files directly from disk. But your Ionic app runs in a browser/WebView environment—this environment doesn’t have access to Node’s fs API, hence the fs.readFileSync is not a function error. On top of that, you were passing just the filename (this.fileToUpload.name) to parseFileItems, which wouldn’t work even in Node unless the file was in the correct directory; in the browser, you need to read the actual file content from the File object instead.
The Solution: Use a Browser-Friendly PDF Library
Instead of pdfreader, use pdf.js—Mozilla’s open-source PDF parsing library designed specifically for browser environments. It can handle the File object directly from your file input without relying on Node.js APIs.
Step 1: Install pdf.js and Its Worker
Run these commands in your project directory to install the required packages:
npm install pdfjs-dist --save npm install pdfjs-dist/build/pdf.worker.min.js --save
Step 2: Update Your TypeScript Code
Replace your existing .ts code with this updated version that uses pdf.js:
import * as pdfjsLib from 'pdfjs-dist'; import pdfWorker from 'pdfjs-dist/build/pdf.worker.min.js'; // Configure the pdf.js worker (required for parsing in browsers) pdfjsLib.GlobalWorkerOptions.workerSrc = pdfWorker; fileToUpload: File; changeListener(files: FileList): void { this.fileToUpload = files.item(0); } async UploadCertificate() { // Check if a file was selected first if (!this.fileToUpload) { console.log('Please select a PDF file before clicking "Get Data"!'); return; } // Read the selected file as an ArrayBuffer const fileReader = new FileReader(); fileReader.onload = async (e) => { const arrayBuffer = e.target.result as ArrayBuffer; try { // Load the PDF document from the ArrayBuffer const pdf = await pdfjsLib.getDocument(arrayBuffer).promise; let fullExtractedText = ''; // Loop through each page to extract text for (let pageNum = 1; pageNum <= pdf.numPages; pageNum++) { const page = await pdf.getPage(pageNum); const textContent = await page.getTextContent(); // Combine all text segments from the page into a single string const pageText = textContent.items.map(item => (item as any).str).join(' '); fullExtractedText += pageText + '\n\n'; } console.log('Full Extracted PDF Text:', fullExtractedText); // Add your logic here to use the extracted text (e.g., display it in the UI) } catch (error) { console.error('Failed to extract PDF text:', error); } }; // Start reading the file content fileReader.readAsArrayBuffer(this.fileToUpload); }
Step 3: Tweak Your HTML for Better Compatibility
Update the accept attribute to use the standard PDF MIME type (your original code works, but this is more reliable across browsers):
<ion-input type="file" accept="application/pdf" (change)="changeListener($event.target.files)"></ion-input> <button ion-button large block type="button" (click)="UploadCertificate()">Get Data</button>
Key Notes
- Async/Await: We use async/await to handle pdf.js’s promise-based operations, making the code cleaner and easier to debug.
- Error Handling: Added checks for missing files and parsing errors to avoid unexpected crashes.
- Worker Setup: The pdf.js worker handles the heavy lifting of PDF parsing in a separate thread, keeping your app responsive.
内容的提问来源于stack exchange,提问作者Nirmalya

