能否使用PDFBox或其他工具检测PDF是否含不可搜索文本及无扫描页?
Detecting Scanned/Unsearchable Pages in PDFs with PDFBox & Tools
Great questions! Let's break down how to address both of your requirements using Apache PDFBox (plus a quick alternative note for Apache Tika):
1. Check if a PDF has NO scanned/unsearchable pages
This means verifying every page in the PDF has extractable, searchable text (images on the page don't matter, per your note). You can do this with PDFBox by extracting text from each page and checking for valid content:
Step 1: Add PDFBox Dependency
First, include the latest stable PDFBox version in your project (Maven example):
<dependency> <groupId>org.apache.pdfbox</groupId> <artifactId>pdfbox</artifactId> <version>2.0.32</version> </dependency>
Step 2: Implementation Code
import org.apache.pdfbox.pdmodel.PDDocument; import org.apache.pdfbox.text.PDFTextStripper; import java.io.File; import java.io.IOException; public class PDFSearchabilityChecker { // Helper method to check if a single page is unsearchable private static boolean isPageUnsearchable(PDDocument document, int pageIndex) throws IOException { PDFTextStripper textStripper = new PDFTextStripper(); // PDFTextStripper uses 1-based page numbering textStripper.setStartPage(pageIndex + 1); textStripper.setEndPage(pageIndex + 1); String pageText = textStripper.getText(document).trim(); // Adjust the threshold based on your needs (e.g., ignore pages with <5 valid characters) return pageText.isEmpty() || pageText.length() < 5; } // Check if the PDF has NO scanned/unsearchable pages (returns true if all pages are searchable) public static boolean hasNoScannedPages(String pdfFilePath) throws IOException { try (PDDocument document = PDDocument.load(new File(pdfFilePath))) { // Handle encrypted PDFs if needed if (document.isEncrypted()) { document.setAllSecurityToBeRemoved(true); } for (int i = 0; i < document.getNumberOfPages(); i++) { if (isPageUnsearchable(document, i)) { return false; // Found an unsearchable page, so scanned pages exist } } return true; // All pages are searchable } } }
2. Detect if a PDF contains ANY unsearchable pages (return boolean)
This is the inverse of the above—we just need to return true as soon as we find one unsearchable page. Add this method to the same class:
// Check if the PDF contains at least one unsearchable page (returns true if any exist) public static boolean containsUnsearchablePages(String pdfFilePath) throws IOException { try (PDDocument document = PDDocument.load(new File(pdfFilePath))) { if (document.isEncrypted()) { document.setAllSecurityToBeRemoved(true); } for (int i = 0; i < document.getNumberOfPages(); i++) { if (isPageUnsearchable(document, i)) { return true; // Found an unsearchable page } } return false; // All pages are searchable } }
Notes & Context
- Threshold Adjustment: The
5-character thresholdinisPageUnsearchableis arbitrary—tweak it based on your use case (e.g., raise it to 10 if you want to ignore pages with trivial noise text). - OCR Edge Case: If a scanned page has been OCR'd and has a hidden text layer, this method will treat it as searchable (which aligns with actual user experience—you can search the text).
- Alternative Tool: Apache Tika: Tika wraps PDFBox and simplifies text extraction, but for granular page-by-page checks, PDFBox is more flexible. If you just need quick full-text extraction, Tika is a good option.
- Encrypted PDFs: The code includes basic handling for encrypted PDFs—adjust if you need to handle password-protected files with specific credentials.
内容的提问来源于stack exchange,提问作者Gamby
相关产品推荐
相关产品推荐

