如何在PDFbox 2.0.27中提取PDF文本与图片并建立关联?
如何提取PDF文本并与对应图片关联?
我需要处理PDF文件中文本与图片的关联需求:PDF中存在描述图片的文本(通常位于图片上方),文件包含多尺寸图片,单页也可能有多组带描述的图片。目前已实现按页码提取文本和图片的功能,但无法将文本与对应的图片关联起来。推测需利用文本和图片的位置元数据,但不知道具体操作方法。当前技术栈为Java 11 + PDFbox 2.0.27,现有代码如下:
public class PdfFileReaderService implements FileReaderService { @Override public String getAllTextFromFile(String filepath) { try (PDDocument document = PDDocument.load(new File(filepath))) { PDFTextStripper stripper = new PDFTextStripper(); return stripper.getText(document); } catch (IOException e) { throw new RuntimeException("Can't get text from file: " + filepath, e); } } @Override public Map<Integer, List<RenderedImage>> getAllImageFromFile(String pathToFile) { Map<Integer, List<RenderedImage>> images = new HashMap<>(); try (PDDocument document = PDDocument.load(new File(pathToFile))) { for (int i = 0; i < document.getNumberOfPages(); i++) { PDPage page = document.getPage(i); List<RenderedImage> pageImages = getImagesFromResources(page.getResources()); images.put(i, pageImages.isEmpty() ? new ArrayList<>() : pageImages); } } catch (IOException e) { throw new RuntimeException("Can't get images from file: " + pathToFile); } return images; } private List<RenderedImage> getImagesFromResources(PDResources resources) throws IOException { List<RenderedImage> images = new ArrayList<>(); for (COSName xObjectName : resources.getXObjectNames()) { PDXObject xObject = resources.getXObject(xObjectName); if (xObject instanceof PDFormXObject) { images.addAll(getImagesFromResources(((PDFormXObject) xObject).getResources())); } else if (xObject instanceof PDImageXObject) { BufferedImage image = ((PDImageXObject) xObject).getImage(); images.add(image); } } return images; } }
核心思路
PDF页面元素(文本、图片)都有精确的坐标信息(PDF坐标系以页面左下角为原点,X轴向右,Y轴向上)。关联文本和图片的关键是获取两者的位置边界,再通过空间位置关系匹配(比如描述文本通常在图片正上方,Y坐标大于图片顶部的Y坐标)。
具体实现步骤
1. 提取带位置信息的文本块
自定义PDFTextStripper,重写writeString方法,捕获每个文本片段的坐标、页码和内容:
// 存储带位置的文本片段 public class TextBlock { private String content; private float x; private float y; private float width; private float height; private int pageNum; // 构造器、getter/setter 省略 } public class PositionAwareTextStripper extends PDFTextStripper { private final List<TextBlock> textBlocks = new ArrayList<>(); private int currentPage; public PositionAwareTextStripper() throws IOException { super(); } @Override protected void startPage(PDPage page) throws IOException { currentPage = getCurrentPageNo(); super.startPage(page); } @Override protected void writeString(String text, List<TextPosition> textPositions) throws IOException { if (text.trim().isEmpty()) { return; } TextPosition firstPos = textPositions.get(0); TextPosition lastPos = textPositions.get(textPositions.size() - 1); TextBlock block = new TextBlock(); block.setContent(text.trim()); block.setX(firstPos.getX()); block.setY(firstPos.getY()); block.setWidth(lastPos.getX() + lastPos.getWidth() - firstPos.getX()); block.setHeight(firstPos.getHeight()); block.setPageNum(currentPage); textBlocks.add(block); super.writeString(text, textPositions); } public List<TextBlock> getTextBlocks() { return textBlocks; } } // 使用方式 public List<TextBlock> getTextWithPositions(String filePath) throws IOException { try (PDDocument document = PDDocument.load(new File(filePath))) { PositionAwareTextStripper stripper = new PositionAwareTextStripper(); stripper.getText(document); return stripper.getTextBlocks(); } }
2. 提取带位置信息的图片
原代码仅提取图片内容,需解析页面内容流获取图片的绘制位置:
// 存储带位置的图片 public class ImageWithPosition { private RenderedImage image; private float x1; // 左下角X坐标 private float y1; // 左下角Y坐标 private float x2; // 右上角X坐标 private float y2; // 右上角Y坐标 private int pageNum; // 构造器、getter/setter 省略 } public class PositionAwareImageExtractor { public List<ImageWithPosition> getImagesWithPositions(String filePath) throws IOException { List<ImageWithPosition> result = new ArrayList<>(); try (PDDocument document = PDDocument.load(new File(filePath))) { for (int pageIdx = 0; pageIdx < document.getNumberOfPages(); pageIdx++) { PDPage page = document.getPage(pageIdx); PDResources resources = page.getResources(); Matrix currentMatrix = new Matrix(); PDStream stream = page.getContents(); if (stream != null) { PDFStreamParser parser = new PDFStreamParser(stream); parser.parse(); List<Object> tokens = parser.getTokens(); for (int i = 0; i < tokens.size(); i++) { Object token = tokens.get(i); if (token instanceof Operator) { Operator op = (Operator) token; String opName = op.getName(); // 处理矩阵变换指令,累积当前坐标变换 if ("cm".equals(opName)) { COSNumber a = (COSNumber) tokens.get(i - 6); COSNumber b = (COSNumber) tokens.get(i - 5); COSNumber c = (COSNumber) tokens.get(i - 4); COSNumber d = (COSNumber) tokens.get(i - 3); COSNumber e = (COSNumber) tokens.get(i - 2); COSNumber f = (COSNumber) tokens.get(i - 1); Matrix matrix = new Matrix(a.floatValue(), b.floatValue(), c.floatValue(), d.floatValue(), e.floatValue(), f.floatValue()); currentMatrix = currentMatrix.multiply(matrix); } // 处理图片绘制指令 if ("Do".equals(opName)) { COSName name = (COSName) tokens.get(i - 1); PDXObject xObject = resources.getXObject(name); if (xObject instanceof PDImageXObject) { PDImageXObject imageXObject = (PDImageXObject) xObject; float imgWidth = imageXObject.getWidth(); float imgHeight = imageXObject.getHeight(); // 计算变换后的实际位置 Point2D.Float bottomLeft = currentMatrix.transformPoint(0, 0); Point2D.Float topRight = currentMatrix.transformPoint(imgWidth, imgHeight); ImageWithPosition imgWithPos = new ImageWithPosition(); imgWithPos.setImage(imageXObject.getImage()); imgWithPos.setX1(bottomLeft.x); imgWithPos.setY1(bottomLeft.y); imgWithPos.setX2(topRight.x); imgWithPos.setY2(topRight.y); imgWithPos.setPageNum(pageIdx + 1); result.add(imgWithPos); } } } } } } } return result; } }
3. 关联文本与图片
根据位置关系匹配,优先选择图片上方、水平范围重叠且距离最近的文本作为描述:
// 存储关联后的结果 public class ImageWithCaption { private RenderedImage image; private String caption; // 构造器、getter/setter 省略 } // 关联逻辑示例 public List<ImageWithCaption> associateTextAndImages(List<TextBlock> textBlocks, List<ImageWithPosition> images) { List<ImageWithCaption> result = new ArrayList<>(); // 按页码分组文本和图片 Map<Integer, List<TextBlock>> textByPage = textBlocks.stream() .collect(Collectors.groupingBy(TextBlock::getPageNum)); Map<Integer, List<ImageWithPosition>> imagesByPage = images.stream() .collect(Collectors.groupingBy(ImageWithPosition::getPageNum)); for (Map.Entry<Integer, List<ImageWithPosition>> entry : imagesByPage.entrySet()) { int pageNum = entry.getKey(); List<TextBlock> pageTexts = textByPage.getOrDefault(pageNum, new ArrayList<>()); List<ImageWithPosition> pageImages = entry.getValue(); for (ImageWithPosition img : pageImages) { // 筛选图片上方、水平重叠的文本,按距离由近到远排序 List<TextBlock> candidates = pageTexts.stream() .filter(t -> t.getY() > img.getY2() && t.getX() + t.getWidth() > img.getX1() && t.getX() < img.getX2()) .sorted(Comparator.comparingDouble(t -> t.getY() - img.getY2())) .collect(Collectors.toList()); ImageWithCaption imgWithCaption = new ImageWithCaption(); imgWithCaption.setImage(img.getImage()); if (!candidates.isEmpty()) { imgWithCaption.setCaption(candidates.get(0).getContent()); } result.add(imgWithCaption); } } return result; }
注意事项
- 坐标系适配:部分PDF会使用裁剪框/媒体框限制显示区域,可通过
page.getCropBox()获取有效区域调整坐标计算。 - 复杂布局处理:如果存在文本环绕图片等复杂布局,需调整匹配逻辑,比如合并相邻文本块、扩大匹配范围。
- 性能优化:解析大PDF时可分页处理,避免内存占用过高。
内容的提问来源于stack exchange,提问作者Vitalii Smahlenko
相关产品推荐
相关产品推荐

