请求获取PDF文本位置的完整代码:解决文本精确坐标获取问题(需包含myLocationExtractionStrategy类)
没问题!我帮你整理了一份基于iText 7的完整Java代码,专门用来提取PDF中文本的精确坐标,其中包含你需要的myLocationExtractionStrategy自定义类,能解决你遇到的文本位置不准确的问题。
依赖准备
首先你需要在项目中引入iText 7的依赖,如果用Maven的话,在pom.xml里添加:
<dependency> <groupId>com.itextpdf</groupId> <artifactId>itext7-core</artifactId> <version>7.2.5</version> </dependency>
如果是Gradle项目,在build.gradle里添加:
implementation 'com.itextpdf:itext7-core:7.2.5'
完整代码实现
import com.itextpdf.kernel.pdf.PdfDocument; import com.itextpdf.kernel.pdf.PdfReader; import com.itextpdf.kernel.pdf.canvas.parser.PdfCanvasProcessor; import com.itextpdf.kernel.pdf.canvas.parser.listener.LocationTextExtractionStrategy; import com.itextpdf.kernel.pdf.canvas.parser.listener.TextRenderInfo; import java.util.ArrayList; import java.util.List; // 自定义文本位置提取策略类 class myLocationExtractionStrategy extends LocationTextExtractionStrategy { // 存储每个文本块的位置和内容信息 private final List<TextLocation> textLocations = new ArrayList<>(); @Override public void renderText(TextRenderInfo renderInfo) { super.renderText(renderInfo); // 获取文本块的左下角坐标(PDF坐标系,原点在左下角) float x = renderInfo.getDescentLine().getStartPoint().get(0); float y = renderInfo.getDescentLine().getStartPoint().get(1); // 获取文本块的右上角坐标 float endX = renderInfo.getAscentLine().getEndPoint().get(0); float endY = renderInfo.getAscentLine().getEndPoint().get(1); // 计算文本块的宽度和高度 float width = endX - x; float height = endY - y; // 获取文本内容 String text = renderInfo.getText(); textLocations.add(new TextLocation(x, y, width, height, text)); } // 获取所有文本位置信息 public List<TextLocation> getTextLocations() { return textLocations; } // 内部类,封装文本位置和内容 public static class TextLocation { private final float x; private final float y; private final float width; private final float height; private final String text; public TextLocation(float x, float y, float width, float height, String text) { this.x = x; this.y = y; this.width = width; this.height = height; this.text = text; } // Getter方法,方便获取各属性 public float getX() { return x; } public float getY() { return y; } public float getWidth() { return width; } public float getHeight() { return height; } public String getText() { return text; } @Override public String toString() { return String.format("文本: '%s' | 坐标(x,y): (%.2f, %.2f) | 宽高: %.2f x %.2f", text, x, y, width, height); } } } // 测试主类 public class PdfTextLocationExtractor { public static void main(String[] args) { String pdfPath = "你的PDF文件路径.pdf"; // 替换成你的PDF路径 try (PdfReader reader = new PdfReader(pdfPath); PdfDocument pdfDoc = new PdfDocument(reader)) { myLocationExtractionStrategy strategy = new myLocationExtractionStrategy(); // 遍历PDF的每一页 for (int pageNum = 1; pageNum <= pdfDoc.getNumberOfPages(); pageNum++) { PdfCanvasProcessor processor = new PdfCanvasProcessor(strategy); processor.processPageContent(pdfDoc.getPage(pageNum)); System.out.println("=== 第" + pageNum + "页文本位置信息 ==="); for (myLocationExtractionStrategy.TextLocation loc : strategy.getTextLocations()) { System.out.println(loc); } // 处理完一页后清空列表,避免跨页数据混淆 strategy.getTextLocations().clear(); } } catch (Exception e) { e.printStackTrace(); } } }
关键说明
myLocationExtractionStrategy继承自iText的LocationTextExtractionStrategy,重写renderText方法来捕获每个文本块的精确坐标- PDF的坐标系原点在左下角,所以y值越大表示位置越靠上
TextRenderInfo还提供了字体、字号、旋转角度等额外信息,如果需要可以在renderText里进一步提取- 代码用
TextLocation内部类封装了文本的位置和内容,方便后续的业务处理
内容的提问来源于stack exchange,提问作者Akash Singh
相关产品推荐
相关产品推荐

