You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用iText7 PDFCleanup无法替换多页PDF所有页面文本的问题

iText PDFCleanup多页文本替换异常问题

场景描述:
我有大量包含SSN等PII数据的多页PDF,分享前需将SSN替换为虚拟值。使用iText及其PDFCleanup插件时遇到异常:可清理任意页面的文本,但仅能替换第一页的文本,后续页面文本无法完成替换。

示例代码:

private void replaceTextInDocument(PdfDocument pdfDocument, String textToReplace, String replacementText) throws IOException {
    CompositeCleanupStrategy strategy = new CompositeCleanupStrategy();
    strategy.add(new RegexBasedCleanupStrategy(textToReplace).setRedactionColor(ColorConstants.WHITE));
    PdfCleaner.autoSweepCleanUp(pdfDocument, strategy);
    System.out.println("----STRATEGY SIZE:  " + strategy.getResultantLocations().size());

    for (IPdfTextLocation location : strategy.getResultantLocations()) {
        System.out.println("----PAGE #: " + location.getPageNumber() + " TEXT: " + location.getText());
        PdfPage page = pdfDocument.getPage(location.getPageNumber() + 1);
        PdfCanvas pdfCanvas = new PdfCanvas(page.newContentStreamAfter(), page.getResources(), page.getDocument());
        Canvas canvas = new Canvas(pdfCanvas, location.getRectangle());
        canvas.add(new Paragraph(replacementText).setFontSize(8).setMarginTop(0f)
              .setRotationAngle(Math.toRadians(90))
        );
    }
}

调试发现:
当待替换文本位于非首页面时,strategy.getResultantLocations()返回集合大小为0,无法进入替换循环,但PDFCleaner本身可正常清理任意页面的文本。

补充背景:
当前处于POC阶段,运行独立Java程序。调试RegexBasedLocationExtractionStrategy的getResultantLocations()方法时,发现parseResults和txt.text始终仅包含首页内容,导致正则匹配失败,返回空集合。目前无法确定DefaultPdfTextLocation中硬编码的0是否为问题根源,寻求以下问题的分析或调试方向:

  • 为何RegexBasedLocationExtractionStrategy无法读取首页以外的内容?

内容的提问来源于stack exchange,提问作者Manchanda. P

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 20:15:29