使用iText7 PDFCleanup无法替换多页PDF所有页面文本的问题
iText PDFCleanup多页文本替换异常问题
场景描述:
我有大量包含SSN等PII数据的多页PDF,分享前需将SSN替换为虚拟值。使用iText及其PDFCleanup插件时遇到异常:可清理任意页面的文本,但仅能替换第一页的文本,后续页面文本无法完成替换。
示例代码:
private void replaceTextInDocument(PdfDocument pdfDocument, String textToReplace, String replacementText) throws IOException { CompositeCleanupStrategy strategy = new CompositeCleanupStrategy(); strategy.add(new RegexBasedCleanupStrategy(textToReplace).setRedactionColor(ColorConstants.WHITE)); PdfCleaner.autoSweepCleanUp(pdfDocument, strategy); System.out.println("----STRATEGY SIZE: " + strategy.getResultantLocations().size()); for (IPdfTextLocation location : strategy.getResultantLocations()) { System.out.println("----PAGE #: " + location.getPageNumber() + " TEXT: " + location.getText()); PdfPage page = pdfDocument.getPage(location.getPageNumber() + 1); PdfCanvas pdfCanvas = new PdfCanvas(page.newContentStreamAfter(), page.getResources(), page.getDocument()); Canvas canvas = new Canvas(pdfCanvas, location.getRectangle()); canvas.add(new Paragraph(replacementText).setFontSize(8).setMarginTop(0f) .setRotationAngle(Math.toRadians(90)) ); } }
调试发现:
当待替换文本位于非首页面时,strategy.getResultantLocations()返回集合大小为0,无法进入替换循环,但PDFCleaner本身可正常清理任意页面的文本。
补充背景:
当前处于POC阶段,运行独立Java程序。调试RegexBasedLocationExtractionStrategy的getResultantLocations()方法时,发现parseResults和txt.text始终仅包含首页内容,导致正则匹配失败,返回空集合。目前无法确定DefaultPdfTextLocation中硬编码的0是否为问题根源,寻求以下问题的分析或调试方向:
- 为何
RegexBasedLocationExtractionStrategy无法读取首页以外的内容?
内容的提问来源于stack exchange,提问作者Manchanda. P
相关产品推荐
相关产品推荐

