You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用文本提取策略?iText7自定义策略下保留PDF表格格式咨询

iText7 提取PDF表格并保留格式的解决方案

要解决PDF表格格式丢失、混合布局表格识别以及空单元格处理的问题,核心是自定义文本提取策略,基于文本片段的坐标信息识别表格结构,而非依赖默认的流式文本提取。以下是具体实现方案:

核心思路

  1. 记录文本片段的精确坐标(x/y轴范围),以此作为表格行列划分的依据;
  2. 区分水平/垂直表格:
    • 水平表格:按Y坐标分组(同一行),按X坐标排序分配列;
    • 垂直表格:按X坐标分组(同一列),按Y坐标排序分配行;
  3. 空单元格直接填充空字符串,保证表格结构完整性。

自定义提取策略实现

import com.itextpdf.kernel.geom.Rectangle;
import com.itextpdf.kernel.pdf.canvas.parser.EventType;
import com.itextpdf.kernel.pdf.canvas.parser.IEventData;
import com.itextpdf.kernel.pdf.canvas.parser.PdfCanvasProcessor;
import com.itextpdf.kernel.pdf.canvas.parser.data.TextRenderInfo;
import com.itextpdf.kernel.pdf.canvas.parser.listener.LocationTextExtractionStrategy;

import java.util.*;

public class TableTextExtractionStrategy extends LocationTextExtractionStrategy {
    private final List<TextChunk> textChunks = new ArrayList<>();

    @Override
    public void eventOccurred(IEventData data, EventType type) {
        if (type == EventType.RENDER_TEXT) {
            TextRenderInfo renderInfo = (TextRenderInfo) data;
            textChunks.add(new TextChunk(renderInfo));
        }
    }

    // 存储文本片段内容与坐标
    private static class TextChunk {
        final String text;
        final float x0, x1, y0, y1;

        public TextChunk(TextRenderInfo renderInfo) {
            this.text = renderInfo.getText();
            Rectangle rect = renderInfo.getAscentLine().getBoundingRectangle();
            this.x0 = rect.getX();
            this.x1 = rect.getX() + rect.getWidth();
            this.y0 = rect.getY();
            this.y1 = rect.getY() + rect.getHeight();
        }
    }

    // 提取表格数据,isHorizontal标记是否为水平布局表格
    public List<List<String>> extractTableData(boolean isHorizontal) {
        List<List<String>> tableData = new ArrayList<>();
        if (textChunks.isEmpty()) return tableData;

        if (isHorizontal) {
            // 按Y坐标分组(容差适配排版误差)
            Map<Float, List<TextChunk>> rowMap = new TreeMap<>(Collections.reverseOrder());
            float tolerance = 2f;
            for (TextChunk chunk : textChunks) {
                float key = Math.round(chunk.getY0() / tolerance) * tolerance;
                rowMap.computeIfAbsent(key, k -> new ArrayList<>()).add(chunk);
            }

            // 获取所有列的X边界
            Set<Float> xBounds = new TreeSet<>();
            rowMap.values().forEach(row -> row.forEach(chunk -> {
                xBounds.add(chunk.getX0());
                xBounds.add(chunk.getX1());
            }));
            List<Float> sortedX = new ArrayList<>(xBounds);

            // 填充每行数据,空单元格留空
            for (List<TextChunk> row : rowMap.values()) {
                List<String> rowData = new ArrayList<>();
                for (int i = 0; i < sortedX.size() - 1; i++) {
                    float colStart = sortedX.get(i);
                    float colEnd = sortedX.get(i + 1);
                    StringBuilder cellText = new StringBuilder();
                    row.forEach(chunk -> {
                        if (chunk.getX0() >= colStart && chunk.getX1() <= colEnd) {
                            cellText.append(chunk.getText());
                        }
                    });
                    rowData.add(cellText.length() == 0 ? "" : cellText.toString());
                }
                tableData.add(rowData);
            }
        } else {
            // 垂直表格处理逻辑:按X分组,按Y划分行
            Map<Float, List<TextChunk>> colMap = new TreeMap<>();
            float tolerance = 2f;
            for (TextChunk chunk : textChunks) {
                float key = Math.round(chunk.getX0() / tolerance) * tolerance;
                colMap.computeIfAbsent(key, k -> new ArrayList<>()).add(chunk);
            }

            Set<Float> yBounds = new TreeSet<>(Collections.reverseOrder());
            colMap.values().forEach(col -> col.forEach(chunk -> {
                yBounds.add(chunk.getY0());
                yBounds.add(chunk.getY1());
            }));
            List<Float> sortedY = new ArrayList<>(yBounds);

            // 构建行数据
            for (int i = 0; i < sortedY.size() - 1; i++) {
                float rowStart = sortedY.get(i);
                float rowEnd = sortedY.get(i + 1);
                List<String> rowData = new ArrayList<>();
                colMap.values().forEach(col -> {
                    StringBuilder cellText = new StringBuilder();
                    col.forEach(chunk -> {
                        if (chunk.getY0() >= rowEnd && chunk.getY1() <= rowStart) {
                            cellText.append(chunk.getText());
                        }
                    });
                    rowData.add(cellText.length() == 0 ? "" : cellText.toString());
                });
                tableData.add(rowData);
            }
        }
        return tableData;
    }
}

使用示例

import com.itextpdf.kernel.pdf.PdfDocument;
import com.itextpdf.kernel.pdf.PdfReader;

import java.io.PrintWriter;
import java.util.List;

public class TableExtractor {
    public static void main(String[] args) throws Exception {
        PdfDocument pdfDoc = new PdfDocument(new PdfReader("input.pdf"));
        TableTextExtractionStrategy strategy = new TableTextExtractionStrategy();
        
        // 处理第一页(假设为水平表格)
        PdfCanvasProcessor processor = new PdfCanvasProcessor(strategy);
        processor.processPageContent(pdfDoc.getPage(1));
        List<List<String>> horizontalTable = strategy.extractTableData(true);
        
        // 输出到文本文件,用制表符分隔列保留结构
        try (PrintWriter writer = new PrintWriter("horizontal_table.txt")) {
            horizontalTable.forEach(row -> writer.println(String.join("\t", row)));
        }

        // 处理第二页(假设为垂直表格)
        strategy = new TableTextExtractionStrategy();
        processor = new PdfCanvasProcessor(strategy);
        processor.processPageContent(pdfDoc.getPage(2));
        List<List<String>> verticalTable = strategy.extractTableData(false);
        
        try (PrintWriter writer = new PrintWriter("vertical_table.txt")) {
            verticalTable.forEach(row -> writer.println(String.join("\t", row)));
        }
        
        pdfDoc.close();
    }
}

关键注意事项

  • 容差调整:tolerance值需根据PDF字体大小、排版精度调整,避免同一行/列的文本被错误分组;
  • 边框识别优化:若PDF包含明确表格边框,可扩展策略解析画布路径事件,通过矩形框精准划分单元格,进一步提升识别准确率;
  • 混合结构处理:若单页同时存在水平/垂直表格,需先划分页面区域,判断每个区域的表格类型后分别处理。

内容的提问来源于stack exchange,提问作者Ibad Ur Rehman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 08:44:58