如何解决React中ReactFileReader无法读取PDF、DOCX文件内容的问题
Hey there! The issue you’re running into is because readAsText() only works for plain text files—PDF and DOCX are binary formats that require specialized parsing libraries to extract their content. Let’s fix this step by step:
Solution: Handle Multiple File Formats in React
First, we’ll add dedicated libraries to parse PDF and DOCX files, then update your component logic to route each file type to the right parser.
Step 1: Install Required Dependencies
Run these commands in your project directory to install the parsing libraries:
npm install pdf-parse mammoth # or use yarn if you prefer yarn add pdf-parse mammoth
pdf-parse: Extracts plain text from PDF filesmammoth: Handles text extraction from DOCX files
Step 2: Update Your Component Code
Here’s the revised DisplayController component with support for all three file types:
import React, { Component } from "react"; import ReactFileReader from 'react-file-reader'; import pdfParse from 'pdf-parse'; import mammoth from 'mammoth'; class DisplayController extends Component { constructor(props){ super(props) this.state = { value: '', fileName: "" } } handleFiles = async files => { const selectedFile = files[0]; const fileExtension = selectedFile.name.split('.').pop().toLowerCase(); try { let extractedContent; // Route file to the correct parser based on extension switch(fileExtension) { case 'txt': extractedContent = await this.readTextFile(selectedFile); break; case 'pdf': extractedContent = await this.readPdfFile(selectedFile); break; case 'docx': extractedContent = await this.readDocxFile(selectedFile); break; default: alert('Unsupported file format!'); return; } alert("Read Data : " + extractedContent); this.setState({ value: extractedContent, fileName: selectedFile.name }); } catch (error) { console.error('Error reading file:', error); alert('Failed to read file. Please try again.'); } } // Helper: Read plain text files readTextFile(file) { return new Promise((resolve, reject) => { const reader = new FileReader(); reader.onload = () => resolve(reader.result); reader.onerror = reject; reader.readAsText(file); }); } // Helper: Parse PDF files async readPdfFile(file) { const arrayBuffer = await this.readFileAsArrayBuffer(file); const pdfResult = await pdfParse(arrayBuffer); return pdfResult.text; } // Helper: Parse DOCX files async readDocxFile(file) { const arrayBuffer = await this.readFileAsArrayBuffer(file); const docxResult = await mammoth.extractRawText({ arrayBuffer }); return docxResult.value; } // Generic helper to read file as ArrayBuffer (used for PDFs/DOCX) readFileAsArrayBuffer(file) { return new Promise((resolve, reject) => { const reader = new FileReader(); reader.onload = () => resolve(reader.result); reader.onerror = reject; reader.readAsArrayBuffer(file); }); } render() { return ( <form> <div className="files"> <ReactFileReader fileTypes={['.pdf','.txt','.docx']} handleFiles={this.handleFiles}> <button className='btn'>Upload</button> </ReactFileReader> </div> </form> ) } } export default DisplayController;
Key Details:
- File Extension Check: We first grab the file’s extension to decide which parsing logic to use.
- Async/Await: Using async/await makes handling the asynchronous file reading and parsing much cleaner.
- Error Handling: Added try/catch blocks to catch and report any issues with file reading or parsing.
- Specialized Parsers: Each file type gets routed to a helper method that uses the appropriate library to extract text.
Extra Tips:
- For large files, consider adding a loading spinner to let users know the file is being processed.
- If you need to retain formatting (like bold text or headings) from DOCX/PDF, check the official docs for
mammothandpdf-parse—they offer additional options for formatted output.
内容的提问来源于stack exchange,提问作者Moh Moh
相关产品推荐
相关产品推荐

