PDF文档是否包含‘不可访问’内容?Java PDF库技术调研
Absolutely—PDFs absolutely can hold content that standard text extraction code (like the PDFBox/iText snippets you shared) fails to capture. Let’s walk through why this happens and how to handle it with Java libraries:
Common Types of "Unreachable" Content
- Scanned/image-based content: If the PDF is just a collection of scanned pages (pixel images, not actual text), tools like PDFBox’s
PDFTextStripperwill return nothing—since there’s no text data to extract. - Hidden text: Text might be set to the same color as the background, or marked with a "hidden" flag in the PDF’s internal structure. Default extraction often skips this.
- Content in annotations/form fields: Standard text extraction usually ignores form input values, sticky notes, or other annotation content.
- Nested/non-standard content: Some PDFs store text inside embedded XObjects, layers, or custom page structures that default extractors don’t traverse.
Fixes for Java PDF Libraries
1. Handling Scanned PDFs (OCR Integration)
For image-only PDFs, you’ll need to combine your PDF library with an OCR tool like Tesseract. Here’s how to do it with PDFBox:
import org.apache.pdfbox.pdmodel.PDDocument; import org.apache.pdfbox.rendering.PDFRenderer; import net.sourceforge.tess4j.Tesseract; import java.awt.image.BufferedImage; import java.io.File; public class OcrPdfExtraction { public static void main(String[] args) throws Exception { PDDocument document = PDDocument.load(new File("scanned-file.pdf")); PDFRenderer renderer = new PDFRenderer(document); Tesseract tesseract = new Tesseract(); StringBuilder extractedText = new StringBuilder(); for (int page = 0; page < document.getNumberOfPages(); page++) { BufferedImage image = renderer.renderImageWithDPI(page, 300); // High DPI for better OCR accuracy extractedText.append(tesseract.doOCR(image)); } System.out.println(extractedText); document.close(); } }
2. Extracting Hidden Text
With PDFBox
Override PDFTextStripper to capture all text, regardless of visibility:
import org.apache.pdfbox.text.PDFTextStripper; import org.apache.pdfbox.text.TextPosition; import java.io.IOException; import java.util.List; public class HiddenTextStripper extends PDFTextStripper { public HiddenTextStripper() throws Exception { super(); } @Override protected void writeString(String text, List<TextPosition> textPositions) throws IOException { // You can add logic here to check if text is truly hidden (e.g., color matching background) // For this example, we'll capture all text including hidden ones super.writeString(text, textPositions); } } // Usage in your existing code: // HiddenTextStripper stripper = new HiddenTextStripper(); // String text = stripper.getText(document);
With iText
Use a custom text extraction strategy to capture all text fragments:
import com.itextpdf.kernel.pdf.PdfDocument; import com.itextpdf.kernel.pdf.PdfReader; import com.itextpdf.kernel.pdf.canvas.parser.PdfCanvasProcessor; import com.itextpdf.kernel.pdf.canvas.parser.listener.LocationTextExtractionStrategy; public class ITextHiddenExtractor { public static void main(String[] args) throws Exception { PdfReader reader = new PdfReader("file.pdf"); PdfDocument pdfDoc = new PdfDocument(reader); LocationTextExtractionStrategy strategy = new LocationTextExtractionStrategy(); for (int page = 1; page <= pdfDoc.getNumberOfPages(); page++) { PdfCanvasProcessor processor = new PdfCanvasProcessor(strategy); processor.processPageContent(pdfDoc.getPage(page)); } System.out.println(strategy.getResultantText()); pdfDoc.close(); reader.close(); } }
(To filter hidden text specifically, extend the strategy to check text rendering modes or color properties against the page background.)
3. Extracting Annotations/Form Fields
With PDFBox
// Extract form fields PDAcroForm acroForm = document.getDocumentCatalog().getAcroForm(); if (acroForm != null) { for (PDField field : acroForm.getFields()) { System.out.println("Field Name: " + field.getFullyQualifiedName() + ", Value: " + field.getValueAsString()); } } // Extract annotations for (PDPage page : document.getPages()) { for (PDAnnotation annotation : page.getAnnotations()) { if (annotation instanceof PDAnnotationTextMarkup) { System.out.println("Annotation Text: " + ((PDAnnotationTextMarkup) annotation).getContents()); } } }
With iText
// Extract form fields PdfAcroForm form = PdfAcroForm.getAcroForm(pdfDoc, false); if (form != null) { form.getFormFields().forEach((name, field) -> System.out.println("Field Name: " + name + ", Value: " + field.getValueAsString()) ); } // Extract annotations for (int page = 1; page <= pdfDoc.getNumberOfPages(); page++) { PdfPage pdfPage = pdfDoc.getPage(page); if (pdfPage.getAnnotations() != null) { pdfPage.getAnnotations().forEach(annotDict -> System.out.println("Annotation Content: " + annotDict.getAsString(PdfName.Contents)) ); } }
4. Handling Nested/Complex Content
For text stored in XObjects or layers, you’ll need to traverse the PDF’s internal structure more deeply:
- PDFBox: Extend
PDFTextStripperto explicitly process embedded XObjects, or usePDPageContentStreamto parse each page’s low-level operations. - iText: Use
PdfCanvasProcessorwith a customEventListenerthat handles XObject invocations, ensuring you process content inside nested objects.
内容的提问来源于stack exchange,提问作者Hector

