You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在PDFbox 2.0.27中提取PDF文本与图片并建立关联?

如何提取PDF文本并与对应图片关联?

我需要处理PDF文件中文本与图片的关联需求:PDF中存在描述图片的文本(通常位于图片上方),文件包含多尺寸图片,单页也可能有多组带描述的图片。目前已实现按页码提取文本和图片的功能,但无法将文本与对应的图片关联起来。推测需利用文本和图片的位置元数据,但不知道具体操作方法。当前技术栈为Java 11 + PDFbox 2.0.27,现有代码如下:

public class PdfFileReaderService implements FileReaderService {
    @Override
    public String getAllTextFromFile(String filepath) {
        try (PDDocument document = PDDocument.load(new File(filepath))) {
            PDFTextStripper stripper = new PDFTextStripper();
            return stripper.getText(document);
        } catch (IOException e) {
            throw new RuntimeException("Can't get text from file: " + filepath, e);
        }
    }

    @Override
    public Map<Integer, List<RenderedImage>> getAllImageFromFile(String pathToFile) {
        Map<Integer, List<RenderedImage>> images = new HashMap<>();
        try (PDDocument document = PDDocument.load(new File(pathToFile))) {
            for (int i = 0; i < document.getNumberOfPages(); i++) {
                PDPage page = document.getPage(i);
                List<RenderedImage> pageImages = getImagesFromResources(page.getResources());
                images.put(i, pageImages.isEmpty() ? new ArrayList<>() : pageImages);
            }
        } catch (IOException e) {
            throw new RuntimeException("Can't get images from file: " + pathToFile);
        }
        return images;
    }


    private List<RenderedImage> getImagesFromResources(PDResources resources) throws IOException {
        List<RenderedImage> images = new ArrayList<>();
        for (COSName xObjectName : resources.getXObjectNames()) {
            PDXObject xObject = resources.getXObject(xObjectName);
            if (xObject instanceof PDFormXObject) {
                images.addAll(getImagesFromResources(((PDFormXObject) xObject).getResources()));
            } else if (xObject instanceof PDImageXObject) {
                BufferedImage image = ((PDImageXObject) xObject).getImage();
                images.add(image);
            }
        }
        return images;
    }
}

核心思路

PDF页面元素(文本、图片)都有精确的坐标信息(PDF坐标系以页面左下角为原点,X轴向右,Y轴向上)。关联文本和图片的关键是获取两者的位置边界,再通过空间位置关系匹配(比如描述文本通常在图片正上方,Y坐标大于图片顶部的Y坐标)。

具体实现步骤

1. 提取带位置信息的文本块

自定义PDFTextStripper,重写writeString方法,捕获每个文本片段的坐标、页码和内容:

// 存储带位置的文本片段
public class TextBlock {
    private String content;
    private float x;
    private float y;
    private float width;
    private float height;
    private int pageNum;

    // 构造器、getter/setter 省略
}

public class PositionAwareTextStripper extends PDFTextStripper {
    private final List<TextBlock> textBlocks = new ArrayList<>();
    private int currentPage;

    public PositionAwareTextStripper() throws IOException {
        super();
    }

    @Override
    protected void startPage(PDPage page) throws IOException {
        currentPage = getCurrentPageNo();
        super.startPage(page);
    }

    @Override
    protected void writeString(String text, List<TextPosition> textPositions) throws IOException {
        if (text.trim().isEmpty()) {
            return;
        }
        TextPosition firstPos = textPositions.get(0);
        TextPosition lastPos = textPositions.get(textPositions.size() - 1);
        
        TextBlock block = new TextBlock();
        block.setContent(text.trim());
        block.setX(firstPos.getX());
        block.setY(firstPos.getY());
        block.setWidth(lastPos.getX() + lastPos.getWidth() - firstPos.getX());
        block.setHeight(firstPos.getHeight());
        block.setPageNum(currentPage);
        
        textBlocks.add(block);
        super.writeString(text, textPositions);
    }

    public List<TextBlock> getTextBlocks() {
        return textBlocks;
    }
}

// 使用方式
public List<TextBlock> getTextWithPositions(String filePath) throws IOException {
    try (PDDocument document = PDDocument.load(new File(filePath))) {
        PositionAwareTextStripper stripper = new PositionAwareTextStripper();
        stripper.getText(document);
        return stripper.getTextBlocks();
    }
}

2. 提取带位置信息的图片

原代码仅提取图片内容,需解析页面内容流获取图片的绘制位置:

// 存储带位置的图片
public class ImageWithPosition {
    private RenderedImage image;
    private float x1; // 左下角X坐标
    private float y1; // 左下角Y坐标
    private float x2; // 右上角X坐标
    private float y2; // 右上角Y坐标
    private int pageNum;

    // 构造器、getter/setter 省略
}

public class PositionAwareImageExtractor {
    public List<ImageWithPosition> getImagesWithPositions(String filePath) throws IOException {
        List<ImageWithPosition> result = new ArrayList<>();
        try (PDDocument document = PDDocument.load(new File(filePath))) {
            for (int pageIdx = 0; pageIdx < document.getNumberOfPages(); pageIdx++) {
                PDPage page = document.getPage(pageIdx);
                PDResources resources = page.getResources();
                Matrix currentMatrix = new Matrix();

                PDStream stream = page.getContents();
                if (stream != null) {
                    PDFStreamParser parser = new PDFStreamParser(stream);
                    parser.parse();
                    List<Object> tokens = parser.getTokens();

                    for (int i = 0; i < tokens.size(); i++) {
                        Object token = tokens.get(i);
                        if (token instanceof Operator) {
                            Operator op = (Operator) token;
                            String opName = op.getName();

                            // 处理矩阵变换指令,累积当前坐标变换
                            if ("cm".equals(opName)) {
                                COSNumber a = (COSNumber) tokens.get(i - 6);
                                COSNumber b = (COSNumber) tokens.get(i - 5);
                                COSNumber c = (COSNumber) tokens.get(i - 4);
                                COSNumber d = (COSNumber) tokens.get(i - 3);
                                COSNumber e = (COSNumber) tokens.get(i - 2);
                                COSNumber f = (COSNumber) tokens.get(i - 1);
                                Matrix matrix = new Matrix(a.floatValue(), b.floatValue(),
                                        c.floatValue(), d.floatValue(),
                                        e.floatValue(), f.floatValue());
                                currentMatrix = currentMatrix.multiply(matrix);
                            }

                            // 处理图片绘制指令
                            if ("Do".equals(opName)) {
                                COSName name = (COSName) tokens.get(i - 1);
                                PDXObject xObject = resources.getXObject(name);
                                if (xObject instanceof PDImageXObject) {
                                    PDImageXObject imageXObject = (PDImageXObject) xObject;
                                    float imgWidth = imageXObject.getWidth();
                                    float imgHeight = imageXObject.getHeight();

                                    // 计算变换后的实际位置
                                    Point2D.Float bottomLeft = currentMatrix.transformPoint(0, 0);
                                    Point2D.Float topRight = currentMatrix.transformPoint(imgWidth, imgHeight);

                                    ImageWithPosition imgWithPos = new ImageWithPosition();
                                    imgWithPos.setImage(imageXObject.getImage());
                                    imgWithPos.setX1(bottomLeft.x);
                                    imgWithPos.setY1(bottomLeft.y);
                                    imgWithPos.setX2(topRight.x);
                                    imgWithPos.setY2(topRight.y);
                                    imgWithPos.setPageNum(pageIdx + 1);

                                    result.add(imgWithPos);
                                }
                            }
                        }
                    }
                }
            }
        }
        return result;
    }
}

3. 关联文本与图片

根据位置关系匹配,优先选择图片上方、水平范围重叠且距离最近的文本作为描述:

// 存储关联后的结果
public class ImageWithCaption {
    private RenderedImage image;
    private String caption;

    // 构造器、getter/setter 省略
}

// 关联逻辑示例
public List<ImageWithCaption> associateTextAndImages(List<TextBlock> textBlocks, List<ImageWithPosition> images) {
    List<ImageWithCaption> result = new ArrayList<>();
    // 按页码分组文本和图片
    Map<Integer, List<TextBlock>> textByPage = textBlocks.stream()
            .collect(Collectors.groupingBy(TextBlock::getPageNum));
    Map<Integer, List<ImageWithPosition>> imagesByPage = images.stream()
            .collect(Collectors.groupingBy(ImageWithPosition::getPageNum));

    for (Map.Entry<Integer, List<ImageWithPosition>> entry : imagesByPage.entrySet()) {
        int pageNum = entry.getKey();
        List<TextBlock> pageTexts = textByPage.getOrDefault(pageNum, new ArrayList<>());
        List<ImageWithPosition> pageImages = entry.getValue();

        for (ImageWithPosition img : pageImages) {
            // 筛选图片上方、水平重叠的文本,按距离由近到远排序
            List<TextBlock> candidates = pageTexts.stream()
                    .filter(t -> t.getY() > img.getY2()
                            && t.getX() + t.getWidth() > img.getX1()
                            && t.getX() < img.getX2())
                    .sorted(Comparator.comparingDouble(t -> t.getY() - img.getY2()))
                    .collect(Collectors.toList());

            ImageWithCaption imgWithCaption = new ImageWithCaption();
            imgWithCaption.setImage(img.getImage());
            if (!candidates.isEmpty()) {
                imgWithCaption.setCaption(candidates.get(0).getContent());
            }
            result.add(imgWithCaption);
        }
    }
    return result;
}

注意事项

  • 坐标系适配:部分PDF会使用裁剪框/媒体框限制显示区域,可通过page.getCropBox()获取有效区域调整坐标计算。
  • 复杂布局处理:如果存在文本环绕图片等复杂布局,需调整匹配逻辑,比如合并相邻文本块、扩大匹配范围。
  • 性能优化:解析大PDF时可分页处理,避免内存占用过高。

内容的提问来源于stack exchange,提问作者Vitalii Smahlenko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 06:54:55