如何使用文本提取策略?iText7自定义策略下保留PDF表格格式咨询
iText7 提取PDF表格并保留格式的解决方案
要解决PDF表格格式丢失、混合布局表格识别以及空单元格处理的问题,核心是自定义文本提取策略,基于文本片段的坐标信息识别表格结构,而非依赖默认的流式文本提取。以下是具体实现方案:
核心思路
- 记录文本片段的精确坐标(x/y轴范围),以此作为表格行列划分的依据;
- 区分水平/垂直表格:
- 水平表格:按Y坐标分组(同一行),按X坐标排序分配列;
- 垂直表格:按X坐标分组(同一列),按Y坐标排序分配行;
- 空单元格直接填充空字符串,保证表格结构完整性。
自定义提取策略实现
import com.itextpdf.kernel.geom.Rectangle; import com.itextpdf.kernel.pdf.canvas.parser.EventType; import com.itextpdf.kernel.pdf.canvas.parser.IEventData; import com.itextpdf.kernel.pdf.canvas.parser.PdfCanvasProcessor; import com.itextpdf.kernel.pdf.canvas.parser.data.TextRenderInfo; import com.itextpdf.kernel.pdf.canvas.parser.listener.LocationTextExtractionStrategy; import java.util.*; public class TableTextExtractionStrategy extends LocationTextExtractionStrategy { private final List<TextChunk> textChunks = new ArrayList<>(); @Override public void eventOccurred(IEventData data, EventType type) { if (type == EventType.RENDER_TEXT) { TextRenderInfo renderInfo = (TextRenderInfo) data; textChunks.add(new TextChunk(renderInfo)); } } // 存储文本片段内容与坐标 private static class TextChunk { final String text; final float x0, x1, y0, y1; public TextChunk(TextRenderInfo renderInfo) { this.text = renderInfo.getText(); Rectangle rect = renderInfo.getAscentLine().getBoundingRectangle(); this.x0 = rect.getX(); this.x1 = rect.getX() + rect.getWidth(); this.y0 = rect.getY(); this.y1 = rect.getY() + rect.getHeight(); } } // 提取表格数据,isHorizontal标记是否为水平布局表格 public List<List<String>> extractTableData(boolean isHorizontal) { List<List<String>> tableData = new ArrayList<>(); if (textChunks.isEmpty()) return tableData; if (isHorizontal) { // 按Y坐标分组(容差适配排版误差) Map<Float, List<TextChunk>> rowMap = new TreeMap<>(Collections.reverseOrder()); float tolerance = 2f; for (TextChunk chunk : textChunks) { float key = Math.round(chunk.getY0() / tolerance) * tolerance; rowMap.computeIfAbsent(key, k -> new ArrayList<>()).add(chunk); } // 获取所有列的X边界 Set<Float> xBounds = new TreeSet<>(); rowMap.values().forEach(row -> row.forEach(chunk -> { xBounds.add(chunk.getX0()); xBounds.add(chunk.getX1()); })); List<Float> sortedX = new ArrayList<>(xBounds); // 填充每行数据,空单元格留空 for (List<TextChunk> row : rowMap.values()) { List<String> rowData = new ArrayList<>(); for (int i = 0; i < sortedX.size() - 1; i++) { float colStart = sortedX.get(i); float colEnd = sortedX.get(i + 1); StringBuilder cellText = new StringBuilder(); row.forEach(chunk -> { if (chunk.getX0() >= colStart && chunk.getX1() <= colEnd) { cellText.append(chunk.getText()); } }); rowData.add(cellText.length() == 0 ? "" : cellText.toString()); } tableData.add(rowData); } } else { // 垂直表格处理逻辑:按X分组,按Y划分行 Map<Float, List<TextChunk>> colMap = new TreeMap<>(); float tolerance = 2f; for (TextChunk chunk : textChunks) { float key = Math.round(chunk.getX0() / tolerance) * tolerance; colMap.computeIfAbsent(key, k -> new ArrayList<>()).add(chunk); } Set<Float> yBounds = new TreeSet<>(Collections.reverseOrder()); colMap.values().forEach(col -> col.forEach(chunk -> { yBounds.add(chunk.getY0()); yBounds.add(chunk.getY1()); })); List<Float> sortedY = new ArrayList<>(yBounds); // 构建行数据 for (int i = 0; i < sortedY.size() - 1; i++) { float rowStart = sortedY.get(i); float rowEnd = sortedY.get(i + 1); List<String> rowData = new ArrayList<>(); colMap.values().forEach(col -> { StringBuilder cellText = new StringBuilder(); col.forEach(chunk -> { if (chunk.getY0() >= rowEnd && chunk.getY1() <= rowStart) { cellText.append(chunk.getText()); } }); rowData.add(cellText.length() == 0 ? "" : cellText.toString()); }); tableData.add(rowData); } } return tableData; } }
使用示例
import com.itextpdf.kernel.pdf.PdfDocument; import com.itextpdf.kernel.pdf.PdfReader; import java.io.PrintWriter; import java.util.List; public class TableExtractor { public static void main(String[] args) throws Exception { PdfDocument pdfDoc = new PdfDocument(new PdfReader("input.pdf")); TableTextExtractionStrategy strategy = new TableTextExtractionStrategy(); // 处理第一页(假设为水平表格) PdfCanvasProcessor processor = new PdfCanvasProcessor(strategy); processor.processPageContent(pdfDoc.getPage(1)); List<List<String>> horizontalTable = strategy.extractTableData(true); // 输出到文本文件,用制表符分隔列保留结构 try (PrintWriter writer = new PrintWriter("horizontal_table.txt")) { horizontalTable.forEach(row -> writer.println(String.join("\t", row))); } // 处理第二页(假设为垂直表格) strategy = new TableTextExtractionStrategy(); processor = new PdfCanvasProcessor(strategy); processor.processPageContent(pdfDoc.getPage(2)); List<List<String>> verticalTable = strategy.extractTableData(false); try (PrintWriter writer = new PrintWriter("vertical_table.txt")) { verticalTable.forEach(row -> writer.println(String.join("\t", row))); } pdfDoc.close(); } }
关键注意事项
- 容差调整:
tolerance值需根据PDF字体大小、排版精度调整,避免同一行/列的文本被错误分组; - 边框识别优化:若PDF包含明确表格边框,可扩展策略解析画布路径事件,通过矩形框精准划分单元格,进一步提升识别准确率;
- 混合结构处理:若单页同时存在水平/垂直表格,需先划分页面区域,判断每个区域的表格类型后分别处理。
内容的提问来源于stack exchange,提问作者Ibad Ur Rehman
相关产品推荐
相关产品推荐

