You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PDFBox中getFontNames返回空时,如何判断PDF是否包含字体?

Detecting Fonts in PDFs with Hidden Resources (e.g., Stamp Annotations)

Great question! The issue you're facing is that page.getResources().getFontNames() only scans resources tied directly to the main page content. Fonts used in elements like your typed "copy" stamp (usually stored as a stamp annotation or form field) live in separate resource dictionaries that your current code doesn't check. That's why PDFBox logs fallback font warnings (it's loading those fonts to render the stamp) but your method returns an empty list.

Here are a few reliable alternatives to getFontNames():

1. Check Annotations and Their Appearance Resources

Stamps and other interactive elements store their visual content in appearance streams, which have their own resource dictionaries. You'll need to iterate through page annotations and inspect these streams:

for (int i = 0; i < pageLimit; ++i) {
    PDPage page = pdf.getPage(i);
    // First check the main page resources (like your original code)
    PDResources pageRes = page.getResources();
    if (!pageRes.getFontNames().isEmpty()) {
        return true;
    }

    // Check all annotations on the page
    for (PDAnnotation annot : page.getAnnotations()) {
        // Handle stamp annotations specifically
        if (annot instanceof PDAnnotationStamp) {
            PDAnnotationStamp stamp = (PDAnnotationStamp) annot;
            PDAppearanceDictionary appearance = stamp.getAppearance();
            if (appearance != null) {
                PDAppearanceEntry normalAppearance = appearance.getNormalAppearance();
                if (normalAppearance != null) {
                    PDAppearanceStream appearanceStream = normalAppearance.getAppearanceStream();
                    if (appearanceStream != null) {
                        PDResources annotRes = appearanceStream.getResources();
                        if (annotRes != null && !annotRes.getFontNames().isEmpty()) {
                            return true;
                        }
                    }
                }
            }
        }

        // Also check form field widgets (if your stamp is part of a form)
        if (annot instanceof PDAnnotationWidget) {
            PDAnnotationWidget widget = (PDAnnotationWidget) annot;
            PDAppearanceDictionary appearance = widget.getAppearance();
            if (appearance != null) {
                PDAppearanceEntry normal = appearance.getNormalAppearance();
                if (normal != null) {
                    PDAppearanceStream stream = normal.getAppearanceStream();
                    if (stream != null) {
                        PDResources res = stream.getResources();
                        if (res != null && !res.getFontNames().isEmpty()) {
                            return true;
                        }
                    }
                }
            }
        }
    }
}

2. Inspect the Document's AcroForm Resources

If your stamp is part of a fillable form, fonts might be stored in the global AcroForm resources or per-field resources:

PDAcroForm acroForm = pdf.getDocumentCatalog().getAcroForm();
if (acroForm != null) {
    // Check global form resources
    PDResources formRes = acroForm.getResources();
    if (formRes != null && !formRes.getFontNames().isEmpty()) {
        return true;
    }

    // Check individual form fields
    for (PDField field : acroForm.getFields()) {
        PDResources fieldRes = field.getResources();
        if (fieldRes != null && !fieldRes.getFontNames().isEmpty()) {
            return true;
        }
    }
}

3. Low-Level Scan for Font Dictionaries

For a foolproof check, you can directly scan all COS objects in the PDF for font dictionaries (any dictionary with Type = Font). This catches fonts anywhere in the document:

boolean hasFont = false;
for (COSBase cosObj : pdf.getDocument().getObjects()) {
    if (cosObj instanceof COSDictionary) {
        COSDictionary dict = (COSDictionary) cosObj;
        if (COSName.FONT.equals(dict.getCOSName(COSName.TYPE))) {
            System.err.println("Found font: " + dict.getNameAsString(COSName.BASE_FONT));
            hasFont = true;
            break;
        }
    }
}
return hasFont;

Key Notes:

  • The fallback font warnings you see confirm that PDFBox does detect the font definitions in the PDF—it just can't load the actual font file, so it uses a system fallback. Your code just wasn't looking in the right place.
  • Pure scanned PDFs (image-only) won't have any font dictionaries at all, so these methods will correctly return false for them.
  • Make sure to handle null checks thoroughly—PDFs can have incomplete or malformed annotations/form fields.

内容的提问来源于stack exchange,提问作者HelloWorld

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:33:44