You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从PDF中提取BufferedImage与ALT文本的键值对(PDFBox/Tika均可)

使用Apache PDFBox提取图像与ALT文本

PDFBox是处理PDF内容的专业工具,能直接访问PDF的结构树和资源,精准关联图像与对应的ALT文本。

1. 添加依赖

如果使用Maven,在pom.xml中添加以下依赖:

<dependency>
    <groupId>org.apache.pdfbox</groupId>
    <artifactId>pdfbox</artifactId>
    <version>2.0.32</version>
</dependency>
<dependency>
    <groupId>org.apache.pdfbox</groupId>
    <artifactId>pdfbox-tools</artifactId>
    <version>2.0.32</version>
</dependency>

2. 代码实现

通过遍历PDF的结构树(StructTree),找到带ALT文本的图像元素,提取图像对象和对应文本:

import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.pdmodel.PDPage;
import org.apache.pdfbox.pdmodel.graphics.image.PDImageXObject;
import org.apache.pdfbox.pdmodel.structure.PDObject;
import org.apache.pdfbox.pdmodel.structure.PDParentNode;
import org.apache.pdfbox.pdmodel.structure.PDStructureElement;
import org.apache.pdfbox.pdmodel.interactive.documentnavigation.destination.PDPageXYZDestination;

import java.awt.image.BufferedImage;
import java.io.File;
import java.io.IOException;
import java.util.HashMap;
import java.util.Map;

public class PdfImageAltExtractor {
    public static Map<BufferedImage, String> extractImageAltPairs(String pdfPath) throws IOException {
        Map<BufferedImage, String> imageAltMap = new HashMap<>();
        try (PDDocument document = PDDocument.load(new File(pdfPath))) {
            var treeRoot = document.getDocumentCatalog().getStructureTreeRoot();
            if (treeRoot == null) {
                return imageAltMap;
            }
            traverseStructElements(treeRoot.getKids(), document, imageAltMap);
        }
        return imageAltMap;
    }

    private static void traverseStructElements(PDParentNode parentNode, PDDocument document, Map<BufferedImage, String> imageAltMap) throws IOException {
        for (PDObject element : parentNode.getKids()) {
            if (!(element instanceof PDStructureElement structElement)) {
                continue;
            }
            // 获取当前结构元素的ALT文本
            String altText = structElement.getAlternateDescription();
            // 找到元素关联的页面
            PDPage page = getAssociatedPage(structElement, document);
            if (page == null || altText == null || altText.isEmpty()) {
                // 递归遍历子元素
                if (structElement.getKids() != null) {
                    traverseStructElements(structElement, document, imageAltMap);
                }
                continue;
            }
            // 遍历页面资源中的图像对象
            page.getResources().getXObjects().forEach((cosName, pdXObject) -> {
                if (pdXObject instanceof PDImageXObject imageXObject) {
                    try {
                        BufferedImage image = imageXObject.getImage();
                        imageAltMap.put(image, altText);
                    } catch (IOException e) {
                        e.printStackTrace();
                    }
                }
            });
            // 递归遍历子元素
            if (structElement.getKids() != null) {
                traverseStructElements(structElement, document, imageAltMap);
            }
        }
    }

    private static PDPage getAssociatedPage(PDStructureElement element, PDDocument document) {
        PDPageXYZDestination dest = (PDPageXYZDestination) element.getPageDestination();
        if (dest != null) {
            return dest.getPage();
        }
        // 若直接获取失败,遍历所有页面查找关联资源
        for (PDPage page : document.getPages()) {
            if (page.getResources().contains(element.getCOSObject())) {
                return page;
            }
        }
        return null;
    }

    public static void main(String[] args) throws IOException {
        String pdfPath = "your/pdf/file/path.pdf";
        Map<BufferedImage, String> result = extractImageAltPairs(pdfPath);
        result.forEach((img, alt) -> System.out.printf("图像ALT文本:%s%n", alt));
    }
}

3. 注意事项

  • 确保目标PDF包含完整的无障碍结构树(StructTree),这是ALT文本存储的核心载体;
  • 部分PDF的ALT文本可能嵌套在子结构元素中,需根据实际文档结构调整遍历逻辑;
  • 若图像是嵌入在内容流而非页面资源中,需额外解析内容流中的图像引用。

使用Apache Tika提取图像与ALT文本

Tika提供了统一的文档解析接口,可将PDF转换为XHTML格式,从中提取图像的ALT文本,再关联图像内容。

1. 添加依赖

Maven依赖如下:

<dependency>
    <groupId>org.apache.tika</groupId>
    <artifactId>tika-core</artifactId>
    <version>2.8.0</version>
</dependency>
<dependency>
    <groupId>org.apache.tika</groupId>
    <artifactId>tika-parsers-standard-package</artifactId>
    <version>2.8.0</version>
</dependency>

2. 代码实现

通过自定义SAX处理器捕获XHTML中的<img>标签(含alt属性),同时使用Tika的嵌入文档提取器获取图像对象:

import org.apache.tika.exception.TikaException;
import org.apache.tika.metadata.Metadata;
import org.apache.tika.parser.ParseContext;
import org.apache.tika.parser.pdf.PDFParser;
import org.apache.tika.sax.BodyContentHandler;
import org.apache.tika.sax.EmbeddedDocumentExtractor;
import org.apache.tika.sax.TeeContentHandler;
import org.xml.sax.Attributes;
import org.xml.sax.SAXException;
import org.xml.sax.helpers.DefaultHandler;

import java.awt.image.BufferedImage;
import java.io.File;
import java.io.FileInputStream;
import java.io.IOException;
import java.io.InputStream;
import java.util.HashMap;
import java.util.Map;
import javax.imageio.ImageIO;

public class TikaImageAltExtractor {
    private static final Map<BufferedImage, String> imageAltMap = new HashMap<>();
    private static String currentAltText;

    public static Map<BufferedImage, String> extractImageAltPairs(String pdfPath) throws IOException, TikaException, SAXException {
        File pdfFile = new File(pdfPath);
        try (InputStream inputStream = new FileInputStream(pdfFile)) {
            Metadata metadata = new Metadata();
            ParseContext context = new ParseContext();

            // 设置自定义嵌入文档提取器,捕获图像
            context.set(EmbeddedDocumentExtractor.class, (parser, stream, docMetadata, outputHtml) -> {
                if ("image".equals(docMetadata.get(Metadata.CONTENT_TYPE))) {
                    try {
                        BufferedImage image = ImageIO.read(stream);
                        if (currentAltText != null && !currentAltText.isEmpty()) {
                            imageAltMap.put(image, currentAltText);
                            currentAltText = null;
                        }
                    } catch (IOException e) {
                        e.printStackTrace();
                    }
                }
            });

            // 自定义SAX处理器,捕获img标签的alt属性
            DefaultHandler altHandler = new DefaultHandler() {
                @Override
                public void startElement(String uri, String localName, String qName, Attributes attributes) throws SAXException {
                    if ("img".equals(qName)) {
                        currentAltText = attributes.getValue("alt");
                    }
                }
            };

            BodyContentHandler contentHandler = new BodyContentHandler();
            TeeContentHandler teeHandler = new TeeContentHandler(contentHandler, altHandler);

            PDFParser parser = new PDFParser();
            parser.parse(inputStream, teeHandler, metadata, context);
        }
        return imageAltMap;
    }

    public static void main(String[] args) throws IOException, TikaException, SAXException {
        String pdfPath = "your/pdf/file/path.pdf";
        Map<BufferedImage, String> result = extractImageAltPairs(pdfPath);
        result.forEach((img, alt) -> System.out.printf("图像ALT文本:%s%n", alt));
    }
}

3. 注意事项

  • Tika将PDF转换为XHTML时,图像的alt属性会直接映射到<img>标签中,需确保解析时捕获该属性;
  • 图像提取依赖EmbeddedDocumentExtractor,需注意图像的内容类型判断;
  • 若PDF中图像与ALT文本的位置对应关系复杂,可能需要结合页码、位置信息做进一步关联。

内容的提问来源于stack exchange,提问作者Tristate

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 17:22:33