You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Electron应用中无需外部服务器的PDF转Txt最优方案咨询

Best Local Ways to Convert PDF to TXT in Electron (No External Servers)

Hey there, I’ve dealt with this exact worker error from pdf2json in Electron before—those worker path issues are a total pain when packaging apps. Let’s break down the best server-free solutions, starting with the most straightforward options.

1. Fix the pdf2json Worker Error (If You Want to Stick With It)

The problem usually stems from Electron’s path restrictions when loading web workers in packaged builds. Here’s how to resolve it:

  • First, when initializing the PDFParser, explicitly specify the worker file path using Node’s require.resolve to avoid path confusion:
    const PDFParser = require("pdf2json");
    const pdfParser = new PDFParser(null, 1);
    
    pdfParser.on("pdfParser_dataError", errData => console.error(errData.parserError));
    pdfParser.on("pdfParser_dataReady", () => {
      const rawText = pdfParser.getRawTextContent();
      // Handle your converted text here
    });
    
    // Critical: Point directly to the worker file
    pdfParser.loadPDF("./your-target.pdf", {
      workerSrc: require.resolve("pdf2json/build/pdf.worker.js")
    });
    
  • When packaging with electron-builder or electron-packager, ensure the worker file gets included. Add this to your package.json if using electron-builder:
    "build": {
      "extraResources": [
        "node_modules/pdf2json/build/pdf.worker.js"
      ]
    }
    

2. Use pdf-parse (Most Reliable & Low-Fuss Option)

If you don’t want to mess with worker configurations, pdf-parse is my go-to. It’s a lightweight wrapper around PDF.js that works entirely in Node.js (no web workers needed), making it perfect for Electron.

Steps:

  1. Install the package:
    npm install pdf-parse
    
  2. Use it in your main or renderer process (just ensure node integration is enabled if using the renderer):
    const fs = require("fs");
    const pdfParse = require("pdf-parse");
    
    async function convertPdfToTxt(pdfFilePath) {
      try {
        const pdfBuffer = fs.readFileSync(pdfFilePath);
        const pdfData = await pdfParse(pdfBuffer);
        return pdfData.text; // This is your extracted plain text
      } catch (error) {
        console.error("PDF conversion failed:", error);
        throw error;
      }
    }
    
    // Example usage
    convertPdfToTxt("./sample.pdf")
      .then(text => {
        console.log("Converted Text:\n", text);
        // Save to file or display to your user
      })
      .catch(err => /* Handle error gracefully */);
    

This library handles all low-level PDF parsing under the hood, and it’s stable in both development and packaged Electron apps.

3. Directly Use pdfjs-dist (For Custom Text Extraction)

If you need more control—like extracting text per page, preserving layout, or targeting specific elements—use Mozilla’s official PDF.js library directly. You’ll need to configure the worker for Node.js, but it’s incredibly flexible.

Steps:

  1. Install the package:
    npm install pdfjs-dist
    
  2. Use it in your main process (renderer process works too with proper configuration):
    const fs = require("fs");
    const pdfjsLib = require("pdfjs-dist/legacy/build/pdf");
    
    // Set the worker source to the Node.js-compatible worker
    pdfjsLib.GlobalWorkerOptions.workerSrc = require.resolve("pdfjs-dist/legacy/build/pdf.worker.js");
    
    async function extractTextPerPage(pdfFilePath) {
      const pdfBuffer = fs.readFileSync(pdfFilePath);
      const pdfDoc = await pdfjsLib.getDocument({ data: pdfBuffer }).promise;
      let fullText = "";
    
      for (let pageNum = 1; pageNum <= pdfDoc.numPages; pageNum++) {
        const page = await pdfDoc.getPage(pageNum);
        const textContent = await page.getTextContent();
        const pageText = textContent.items.map(item => item.str).join(" ");
        fullText += `Page ${pageNum}:\n${pageText}\n\n`;
      }
    
      return fullText;
    }
    
    // Example usage
    extractTextPerPage("./sample.pdf").then(text => console.log(text));
    

Which One Should You Choose?

  • Stick with pdf2json: Only if you already have code using it and want to avoid rewriting.
  • Go with pdf-parse: Best for most use cases—simple, reliable, no worker headaches.
  • Use pdfjs-dist: If you need custom text extraction logic (e.g., page-by-page output, formatting preservation).

内容的提问来源于stack exchange,提问作者Israel Zebulon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:47:23