iText提取PDF文本出现异常字符及行判定错误问题
解决iText提取PDF文本时的异常空格与行判定问题
问题根源
- 异常空格:iText默认空格判定阈值过高,将PDF中同一单词内的微小字符偏移误判为空格;行首空格则是文本块起始坐标与行首基准的偏差被识别为前导空格。
- 行判定错误:默认行间距阈值过大,把相邻但分属不同行的文本块归为同一行。
自定义提取策略解决
通过继承LocationTextExtractionStrategy,重写关键方法调整判定逻辑:
import com.itextpdf.kernel.geom.Vector; import com.itextpdf.kernel.pdf.canvas.parser.data.TextChunk; import com.itextpdf.kernel.pdf.canvas.parser.listener.LocationTextExtractionStrategy; import java.util.ArrayList; import java.util.List; public class CustomTextExtractionStrategy extends LocationTextExtractionStrategy { // 自定义空格判定阈值(根据PDF实际排版调整,示例设为2.0f) private static final float SPACE_THRESHOLD = 2.0f; // 自定义行间距阈值(区分不同行的Y坐标差,示例设为8.0f) private static final float LINE_SPACING_THRESHOLD = 8.0f; @Override protected void processTextChunk(TextChunk chunk) { List<TextChunk> chunks = getChunks(); if (chunks.isEmpty()) { chunks.add(chunk); return; } TextChunk lastChunk = chunks.get(chunks.size() - 1); float distance = chunk.getLocation().distanceFromEndOf(lastChunk.getLocation()); // 仅当间距超过阈值且不属于同一单词时,才插入空格 if (distance > SPACE_THRESHOLD && !isSameWord(lastChunk, chunk)) { chunks.add(new TextChunk(" ", lastChunk.getLocation())); } chunks.add(chunk); } // 判断两个文本块是否属于同一单词 private boolean isSameWord(TextChunk a, TextChunk b) { Vector aStart = a.getLocation().getStartLocation(); Vector bStart = b.getLocation().getStartLocation(); // 同一字体、Y坐标偏差极小、X偏移未超过空格阈值 return a.getFont().equals(b.getFont()) && Math.abs(aStart.getY() - bStart.getY()) < 1.0f && (bStart.getX() - a.getLocation().getEndLocation().getX()) < SPACE_THRESHOLD; } @Override protected List<TextChunk> filterTextChunks(List<TextChunk> chunks, boolean hasOverlap) { List<TextChunk> filteredChunks = new ArrayList<>(); TextChunk previousChunk = null; for (TextChunk chunk : chunks) { if (previousChunk != null) { float yDiff = Math.abs(chunk.getLocation().getStartLocation().getY() - previousChunk.getLocation().getStartLocation().getY()); // Y坐标差超过阈值则插入换行符 if (yDiff > LINE_SPACING_THRESHOLD) { filteredChunks.add(new TextChunk("\n", previousChunk.getLocation())); } } filteredChunks.add(chunk); previousChunk = chunk; } return super.filterTextChunks(filteredChunks, hasOverlap); } }
调用自定义策略
替换默认策略,使用自定义类提取文本:
import com.itextpdf.kernel.pdf.PdfDocument; import com.itextpdf.kernel.pdf.PdfReader; import com.itextpdf.kernel.pdf.canvas.parser.PdfTextExtractor; import java.io.IOException; public class PdfTextExtractorDemo { public static void main(String[] args) throws IOException { PdfReader reader = new PdfReader("target-file.pdf"); PdfDocument pdfDoc = new PdfDocument(reader); StringBuilder extractedText = new StringBuilder(); for (int pageNum = 1; pageNum <= pdfDoc.getNumberOfPages(); pageNum++) { CustomTextExtractionStrategy strategy = new CustomTextExtractionStrategy(); String pageText = PdfTextExtractor.getTextFromPage(pdfDoc.getPage(pageNum), strategy); extractedText.append(pageText).append("\n"); } pdfDoc.close(); reader.close(); System.out.println(extractedText.toString()); } }
针对PDF流片段的额外调整
结合你提供的PDF流片段,补充两个优化点:
- 行首异常空格:在
processTextChunk方法中,若当前是行首第一个文本块,直接跳过前导空格判定,避免将起始坐标偏差误判为空格。 - 特定单词的偏移处理:如果“lillyce”是固定出现的单词,可在
isSameWord方法中加入针对性判断,强制将该单词的拆分块合并。
内容的提问来源于stack exchange,提问作者vishal bhardwaj
相关产品推荐
相关产品推荐

