You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PDF文档是否包含‘不可访问’内容?Java PDF库技术调研

Can PDFs Contain "Unreachable" Content That Regular Extraction Tools Miss?

Absolutely—PDFs absolutely can hold content that standard text extraction code (like the PDFBox/iText snippets you shared) fails to capture. Let’s walk through why this happens and how to handle it with Java libraries:

Common Types of "Unreachable" Content

  • Scanned/image-based content: If the PDF is just a collection of scanned pages (pixel images, not actual text), tools like PDFBox’s PDFTextStripper will return nothing—since there’s no text data to extract.
  • Hidden text: Text might be set to the same color as the background, or marked with a "hidden" flag in the PDF’s internal structure. Default extraction often skips this.
  • Content in annotations/form fields: Standard text extraction usually ignores form input values, sticky notes, or other annotation content.
  • Nested/non-standard content: Some PDFs store text inside embedded XObjects, layers, or custom page structures that default extractors don’t traverse.

Fixes for Java PDF Libraries

1. Handling Scanned PDFs (OCR Integration)

For image-only PDFs, you’ll need to combine your PDF library with an OCR tool like Tesseract. Here’s how to do it with PDFBox:

import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.rendering.PDFRenderer;
import net.sourceforge.tess4j.Tesseract;
import java.awt.image.BufferedImage;
import java.io.File;

public class OcrPdfExtraction {
    public static void main(String[] args) throws Exception {
        PDDocument document = PDDocument.load(new File("scanned-file.pdf"));
        PDFRenderer renderer = new PDFRenderer(document);
        Tesseract tesseract = new Tesseract();
        
        StringBuilder extractedText = new StringBuilder();
        for (int page = 0; page < document.getNumberOfPages(); page++) {
            BufferedImage image = renderer.renderImageWithDPI(page, 300); // High DPI for better OCR accuracy
            extractedText.append(tesseract.doOCR(image));
        }
        
        System.out.println(extractedText);
        document.close();
    }
}

2. Extracting Hidden Text

With PDFBox

Override PDFTextStripper to capture all text, regardless of visibility:

import org.apache.pdfbox.text.PDFTextStripper;
import org.apache.pdfbox.text.TextPosition;
import java.io.IOException;
import java.util.List;

public class HiddenTextStripper extends PDFTextStripper {
    public HiddenTextStripper() throws Exception {
        super();
    }

    @Override
    protected void writeString(String text, List<TextPosition> textPositions) throws IOException {
        // You can add logic here to check if text is truly hidden (e.g., color matching background)
        // For this example, we'll capture all text including hidden ones
        super.writeString(text, textPositions);
    }
}

// Usage in your existing code:
// HiddenTextStripper stripper = new HiddenTextStripper();
// String text = stripper.getText(document);

With iText

Use a custom text extraction strategy to capture all text fragments:

import com.itextpdf.kernel.pdf.PdfDocument;
import com.itextpdf.kernel.pdf.PdfReader;
import com.itextpdf.kernel.pdf.canvas.parser.PdfCanvasProcessor;
import com.itextpdf.kernel.pdf.canvas.parser.listener.LocationTextExtractionStrategy;

public class ITextHiddenExtractor {
    public static void main(String[] args) throws Exception {
        PdfReader reader = new PdfReader("file.pdf");
        PdfDocument pdfDoc = new PdfDocument(reader);
        
        LocationTextExtractionStrategy strategy = new LocationTextExtractionStrategy();
        for (int page = 1; page <= pdfDoc.getNumberOfPages(); page++) {
            PdfCanvasProcessor processor = new PdfCanvasProcessor(strategy);
            processor.processPageContent(pdfDoc.getPage(page));
        }
        
        System.out.println(strategy.getResultantText());
        pdfDoc.close();
        reader.close();
    }
}

(To filter hidden text specifically, extend the strategy to check text rendering modes or color properties against the page background.)

3. Extracting Annotations/Form Fields

With PDFBox

// Extract form fields
PDAcroForm acroForm = document.getDocumentCatalog().getAcroForm();
if (acroForm != null) {
    for (PDField field : acroForm.getFields()) {
        System.out.println("Field Name: " + field.getFullyQualifiedName() + ", Value: " + field.getValueAsString());
    }
}

// Extract annotations
for (PDPage page : document.getPages()) {
    for (PDAnnotation annotation : page.getAnnotations()) {
        if (annotation instanceof PDAnnotationTextMarkup) {
            System.out.println("Annotation Text: " + ((PDAnnotationTextMarkup) annotation).getContents());
        }
    }
}

With iText

// Extract form fields
PdfAcroForm form = PdfAcroForm.getAcroForm(pdfDoc, false);
if (form != null) {
    form.getFormFields().forEach((name, field) -> 
        System.out.println("Field Name: " + name + ", Value: " + field.getValueAsString())
    );
}

// Extract annotations
for (int page = 1; page <= pdfDoc.getNumberOfPages(); page++) {
    PdfPage pdfPage = pdfDoc.getPage(page);
    if (pdfPage.getAnnotations() != null) {
        pdfPage.getAnnotations().forEach(annotDict -> 
            System.out.println("Annotation Content: " + annotDict.getAsString(PdfName.Contents))
        );
    }
}

4. Handling Nested/Complex Content

For text stored in XObjects or layers, you’ll need to traverse the PDF’s internal structure more deeply:

  • PDFBox: Extend PDFTextStripper to explicitly process embedded XObjects, or use PDPageContentStream to parse each page’s low-level operations.
  • iText: Use PdfCanvasProcessor with a custom EventListener that handles XObject invocations, ensuring you process content inside nested objects.

内容的提问来源于stack exchange,提问作者Hector

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:03:13