You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Java+PDFBox提取PDF文本时出现错误,请求技术协助

Troubleshooting PDF Text Extraction with Apache PDFBox

Hey there! Let's work through your PDF text extraction issue with PDFBox. I notice a couple of areas we can tweak in your setup and code to get things running smoothly.

First: Update Your PDFBox Dependency

You're using version 2.0.4, which is quite outdated (released back in 2017). Newer versions fix tons of bugs and add better compatibility with modern PDFs. Let's upgrade to a recent stable version—here's what your Maven dependency should look like:

<dependencies>
    <dependency>
        <groupId>org.apache.pdfbox</groupId>
        <artifactId>pdfbox</artifactId>
        <version>2.0.33</version> <!-- Use the latest stable version available -->
    </dependency>
</dependencies>

Fix Your Extraction Code

Your code snippet cuts off, but I can see you're using older classes like PDFParser and COSDocument—these are deprecated in PDFBox 2.x. The modern, simplified way to extract text uses PDDocument directly along with PDFTextStripper. Here's a complete, working example:

import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.text.PDFTextStripper;
import java.io.File;
import java.io.IOException;

public class PDFTextExtractor {
    public static void main(String[] args) {
        String pdfFilePath = "path/to/your/target.pdf";
        
        try (PDDocument document = PDDocument.load(new File(pdfFilePath))) {
            // Handle encrypted PDFs first
            if (document.isEncrypted()) {
                // Try empty password first (common for view-only protected PDFs)
                document.decrypt("");
            }
            
            PDFTextStripper textStripper = new PDFTextStripper();
            String extractedText = textStripper.getText(document);
            
            System.out.println("Extracted Text:\n" + extractedText);
        } catch (IOException e) {
            e.printStackTrace();
            System.err.println("Error extracting text: " + e.getMessage());
        }
    }
}

Common Issues to Check If You Still Get Errors

If you're still seeing problems after updating, here are some key things to verify:

  • Corrupted PDF File: Test with a simple, known-working PDF (like a sample from Apache's official docs) to rule out a damaged file.
  • File Permissions: Make sure your application has read access to the PDF file's path.
  • Strongly Encrypted PDFs: If the empty password doesn't work, you'll need the correct user password to decrypt the document.
  • Dependency Conflicts: Check your build file for other PDF-related libraries (like iText) that might clash with PDFBox.

内容的提问来源于stack exchange,提问作者Sveta Tulova

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:02:44