使用Java+PDFBox提取PDF文本时出现错误,请求技术协助
Hey there! Let's work through your PDF text extraction issue with PDFBox. I notice a couple of areas we can tweak in your setup and code to get things running smoothly.
First: Update Your PDFBox Dependency
You're using version 2.0.4, which is quite outdated (released back in 2017). Newer versions fix tons of bugs and add better compatibility with modern PDFs. Let's upgrade to a recent stable version—here's what your Maven dependency should look like:
<dependencies> <dependency> <groupId>org.apache.pdfbox</groupId> <artifactId>pdfbox</artifactId> <version>2.0.33</version> <!-- Use the latest stable version available --> </dependency> </dependencies>
Fix Your Extraction Code
Your code snippet cuts off, but I can see you're using older classes like PDFParser and COSDocument—these are deprecated in PDFBox 2.x. The modern, simplified way to extract text uses PDDocument directly along with PDFTextStripper. Here's a complete, working example:
import org.apache.pdfbox.pdmodel.PDDocument; import org.apache.pdfbox.text.PDFTextStripper; import java.io.File; import java.io.IOException; public class PDFTextExtractor { public static void main(String[] args) { String pdfFilePath = "path/to/your/target.pdf"; try (PDDocument document = PDDocument.load(new File(pdfFilePath))) { // Handle encrypted PDFs first if (document.isEncrypted()) { // Try empty password first (common for view-only protected PDFs) document.decrypt(""); } PDFTextStripper textStripper = new PDFTextStripper(); String extractedText = textStripper.getText(document); System.out.println("Extracted Text:\n" + extractedText); } catch (IOException e) { e.printStackTrace(); System.err.println("Error extracting text: " + e.getMessage()); } } }
Common Issues to Check If You Still Get Errors
If you're still seeing problems after updating, here are some key things to verify:
- Corrupted PDF File: Test with a simple, known-working PDF (like a sample from Apache's official docs) to rule out a damaged file.
- File Permissions: Make sure your application has read access to the PDF file's path.
- Strongly Encrypted PDFs: If the empty password doesn't work, you'll need the correct user password to decrypt the document.
- Dependency Conflicts: Check your build file for other PDF-related libraries (like iText) that might clash with PDFBox.
内容的提问来源于stack exchange,提问作者Sveta Tulova

