Node.js如何用单一模块读取更多文件类型?textract无法读取部分PDF与RTF
First off, I feel your pain with textract—great for some formats but falls short on edge-case PDFs and RTF. The best single-module solution I’ve found that covers almost all common (and even some niche) file types is Apache Tika's Node.js binding (node-tika). It’s built on the robust Apache Tika library, which supports over 1000 file formats, including the ones you’re struggling with.
How to Use node-tika
1. Install the Package
First, pull it in via npm:
npm install node-tika
2. Basic Usage Example
Here’s a straightforward script that handles PDF, RTF, and other common file types effortlessly:
const Tika = require('node-tika'); const fs = require('fs').promises; async function extractTextFromFile(filePath) { try { const fileBuffer = await fs.readFile(filePath); const result = await Tika.extractText(fileBuffer); return result.text; } catch (error) { console.error(`Failed to extract text from ${filePath}:`, error); throw error; } } // Test with your files (async () => { const pdfContent = await extractTextFromFile('./troublesome.pdf'); console.log('PDF Content:\n', pdfContent); const rtfContent = await extractTextFromFile('./sample.rtf'); console.log('\nRTF Content:\n', rtfContent); const docxContent = await extractTextFromFile('./document.docx'); console.log('\nDOCX Content:\n', docxContent); })();
3. Handling Edge-Case PDFs
Unlike textract, node-tika works with encrypted PDFs (just pass the password as an option) and even scanned PDFs if you have Tesseract OCR set up (it integrates seamlessly for OCR support). Here’s how to tackle encrypted files:
const result = await Tika.extractText(fileBuffer, { password: 'your-pdf-password' });
Alternative: Lighter Weight Option
If node-tika feels too heavy (it bundles the Tika JAR under the hood), try @extractus/text-extractor—a modern, promise-based module that supports all the formats you need without the Java dependency. Install it like this:
npm install @extractus/text-extractor
And use it with minimal code:
const { extract } = require('@extractus/text-extractor'); async function getFileContent(filePath) { const content = await extract(filePath); return content; }
Key Takeaways
- Both modules support way more formats than textract: RTF, PDF (encrypted included), DOC, DOCX, XLS, XLSX, PPT, PPTX, CSV, Markdown, and even image files with OCR.
- For scanned PDFs, ensure Tesseract is installed on your system—both modules will alert you if it’s missing.
内容的提问来源于stack exchange,提问作者ATUL SHARMA

