You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

iText提取PDF文本出现异常字符及行判定错误问题

解决iText提取PDF文本时的异常空格与行判定问题

问题根源

  1. 异常空格:iText默认空格判定阈值过高,将PDF中同一单词内的微小字符偏移误判为空格;行首空格则是文本块起始坐标与行首基准的偏差被识别为前导空格。
  2. 行判定错误:默认行间距阈值过大,把相邻但分属不同行的文本块归为同一行。

自定义提取策略解决

通过继承LocationTextExtractionStrategy,重写关键方法调整判定逻辑:

import com.itextpdf.kernel.geom.Vector;
import com.itextpdf.kernel.pdf.canvas.parser.data.TextChunk;
import com.itextpdf.kernel.pdf.canvas.parser.listener.LocationTextExtractionStrategy;

import java.util.ArrayList;
import java.util.List;

public class CustomTextExtractionStrategy extends LocationTextExtractionStrategy {
    // 自定义空格判定阈值(根据PDF实际排版调整,示例设为2.0f)
    private static final float SPACE_THRESHOLD = 2.0f;
    // 自定义行间距阈值(区分不同行的Y坐标差,示例设为8.0f)
    private static final float LINE_SPACING_THRESHOLD = 8.0f;

    @Override
    protected void processTextChunk(TextChunk chunk) {
        List<TextChunk> chunks = getChunks();
        if (chunks.isEmpty()) {
            chunks.add(chunk);
            return;
        }

        TextChunk lastChunk = chunks.get(chunks.size() - 1);
        float distance = chunk.getLocation().distanceFromEndOf(lastChunk.getLocation());

        // 仅当间距超过阈值且不属于同一单词时,才插入空格
        if (distance > SPACE_THRESHOLD && !isSameWord(lastChunk, chunk)) {
            chunks.add(new TextChunk(" ", lastChunk.getLocation()));
        }
        chunks.add(chunk);
    }

    // 判断两个文本块是否属于同一单词
    private boolean isSameWord(TextChunk a, TextChunk b) {
        Vector aStart = a.getLocation().getStartLocation();
        Vector bStart = b.getLocation().getStartLocation();
        // 同一字体、Y坐标偏差极小、X偏移未超过空格阈值
        return a.getFont().equals(b.getFont())
                && Math.abs(aStart.getY() - bStart.getY()) < 1.0f
                && (bStart.getX() - a.getLocation().getEndLocation().getX()) < SPACE_THRESHOLD;
    }

    @Override
    protected List<TextChunk> filterTextChunks(List<TextChunk> chunks, boolean hasOverlap) {
        List<TextChunk> filteredChunks = new ArrayList<>();
        TextChunk previousChunk = null;

        for (TextChunk chunk : chunks) {
            if (previousChunk != null) {
                float yDiff = Math.abs(chunk.getLocation().getStartLocation().getY()
                        - previousChunk.getLocation().getStartLocation().getY());
                // Y坐标差超过阈值则插入换行符
                if (yDiff > LINE_SPACING_THRESHOLD) {
                    filteredChunks.add(new TextChunk("\n", previousChunk.getLocation()));
                }
            }
            filteredChunks.add(chunk);
            previousChunk = chunk;
        }
        return super.filterTextChunks(filteredChunks, hasOverlap);
    }
}

调用自定义策略

替换默认策略,使用自定义类提取文本:

import com.itextpdf.kernel.pdf.PdfDocument;
import com.itextpdf.kernel.pdf.PdfReader;
import com.itextpdf.kernel.pdf.canvas.parser.PdfTextExtractor;

import java.io.IOException;

public class PdfTextExtractorDemo {
    public static void main(String[] args) throws IOException {
        PdfReader reader = new PdfReader("target-file.pdf");
        PdfDocument pdfDoc = new PdfDocument(reader);
        StringBuilder extractedText = new StringBuilder();

        for (int pageNum = 1; pageNum <= pdfDoc.getNumberOfPages(); pageNum++) {
            CustomTextExtractionStrategy strategy = new CustomTextExtractionStrategy();
            String pageText = PdfTextExtractor.getTextFromPage(pdfDoc.getPage(pageNum), strategy);
            extractedText.append(pageText).append("\n");
        }

        pdfDoc.close();
        reader.close();
        System.out.println(extractedText.toString());
    }
}

针对PDF流片段的额外调整

结合你提供的PDF流片段,补充两个优化点:

  1. 行首异常空格:在processTextChunk方法中,若当前是行首第一个文本块,直接跳过前导空格判定,避免将起始坐标偏差误判为空格。
  2. 特定单词的偏移处理:如果“lillyce”是固定出现的单词,可在isSameWord方法中加入针对性判断,强制将该单词的拆分块合并。

内容的提问来源于stack exchange,提问作者vishal bhardwaj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 17:25:32