You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用PDFBox移除PDF中的{{place_holder}}占位符标签

问题描述

我有一个包含{{place_holder}}格式占位符标签的PDF文档,希望使用PDFBox库移除文档中所有此类标签。

我尝试过一段示例代码但无效,代码如下:

@Override
protected void write(ContentStreamWriter contentStreamWriter, Operator operator, List<COSBase> operands) throws IOException {
    String recentText = recentChars.toString();
    recentChars.setLength(0);

    String operatorString = operator.getName();

    if (TEXT_SHOWING_OPERATORS.contains(operatorString) && "{{full_name}}".equals(recentText))
    {
        return;
    }

    super.write(contentStreamWriter, operator, operands);
}

更新1

以下代码对部分PDF文件有效,但对Microsoft Word导出的PDF无效。由于我的场景中PDF可能由任意系统生成,因此需要一个更通用的解决方案。

有效但不通用的代码:

PdfContentStreamEditor editor = new PdfContentStreamEditor(document, page) {
    @Override
    protected void write(ContentStreamWriter contentStreamWriter, Operator operator, List<COSBase> operands) throws IOException {
        String operatorString = operator.getName();
        if (TEXT_SHOWING_OPERATORS.contains(operatorString))
        {
            if(operands.get(0) instanceof COSString ){
                COSString str= (COSString) operands.get(0);
                String text=str.getString();
                String updated= extractStringsBetweenCurlyBraces(text);
                if(!text.equals(updated)){
                    str.setValue(updated.getBytes());
                }
            }
            if(operands.get(0) instanceof COSArray ){
                Iterator var7 =  ((COSArray) operands.get(0)).iterator();
                while(var7.hasNext()) {
                    COSBase obj = (COSBase) var7.next();
                    if (obj instanceof COSString) {
                        COSString str= (COSString) obj;
                        String text=str.getString();
                        String updated= extractStringsBetweenCurlyBraces(text);
                        str.setValue(updated.getBytes());
                    }
                }
            }
        }
        super.write(contentStreamWriter, operator, operands);
    }
    final List<String> TEXT_SHOWING_OPERATORS = Arrays.asList("Tj", "'", "\"", "TJ");
};
editor.processPage(page);



public static String extractStringsBetweenCurlyBraces(String input) {
    Pattern pattern = Pattern.compile("\\{\\{[^}]*\\}\\}||\\{\\{.*$");
    Matcher matcher = pattern.matcher(input);
    while (matcher.find()) {
        String match = matcher.group();
        String replacement = " ".repeat(match.length()+7);
        input= input.replace(match,replacement);
    }

    pattern = Pattern.compile("^.*?\\}\\}");
    matcher = pattern.matcher(input);
    while (matcher.find()) {
        String match = matcher.group();
        String replacement = " ".repeat(match.length()+7);
        input= input.replace(match,replacement);
    }
    return input;
}

通用解决方案

要处理任意系统生成的PDF,核心问题是占位符可能被拆分到多个文本绘制操作中,或者单个文本操作内的多个字符串片段里。我们需要跟踪当前累积的文本,识别完整的{{...}}占位符,再决定是否跳过对应的绘制操作。

实现代码

import org.apache.pdfbox.contentstream.operator.Operator;
import org.apache.pdfbox.cos.COSBase;
import org.apache.pdfbox.cos.COSString;
import org.apache.pdfbox.pdfwriter.ContentStreamWriter;
import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.pdmodel.PDPage;
import org.apache.pdfbox.util.PDFStreamEngine;

import java.io.IOException;
import java.util.List;
import java.util.regex.Matcher;
import java.util.regex.Pattern;

public class PlaceholderRemover {
    private static final List<String> TEXT_SHOWING_OPERATORS = List.of("Tj", "'", "\"", "TJ");
    private static final Pattern PLACEHOLDER_PATTERN = Pattern.compile("\\{\\{.*?\\}\\}");

    public static void removePlaceholders(PDDocument document, PDPage page) throws IOException {
        new PdfContentStreamEditor(document, page) {
            private final StringBuilder currentTextBuffer = new StringBuilder();
            private boolean inPlaceholder = false;

            @Override
            protected void write(ContentStreamWriter writer, Operator operator, List<COSBase> operands) throws IOException {
                String opName = operator.getName();
                if (TEXT_SHOWING_OPERATORS.contains(opName)) {
                    StringBuilder opText = new StringBuilder();
                    extractTextFromOperands(operands, opText);

                    currentTextBuffer.append(opText);
                    String fullText = currentTextBuffer.toString();
                    Matcher matcher = PLACEHOLDER_PATTERN.matcher(fullText);

                    if (matcher.find()) {
                        // 替换占位符为等长空格,保留布局
                        String cleanedText = fullText.replaceAll("\\{\\{.*?\\}\\} ", " ".repeat(matcher.group().length()));
                        updateOperandsWithCleanedText(operands, cleanedText);
                        currentTextBuffer.setLength(0);
                        super.write(writer, operator, operands);
                    } else if (fullText.startsWith("{{")) {
                        inPlaceholder = true;
                    } else if (fullText.endsWith("}}") && inPlaceholder) {
                        currentTextBuffer.setLength(0);
                        inPlaceholder = false;
                    } else {
                        super.write(writer, operator, operands);
                        currentTextBuffer.setLength(0);
                    }
                } else {
                    currentTextBuffer.setLength(0);
                    inPlaceholder = false;
                    super.write(writer, operator, operands);
                }
            }

            private void extractTextFromOperands(List<COSBase> operands, StringBuilder textBuilder) {
                if (operands.get(0) instanceof COSString cosString) {
                    textBuilder.append(cosString.getString());
                } else if (operands.get(0) instanceof COSArray cosArray) {
                    for (COSBase base : cosArray) {
                        if (base instanceof COSString str) {
                            textBuilder.append(str.getString());
                        }
                    }
                }
            }

            private void updateOperandsWithCleanedText(List<COSBase> operands, String cleanedText) throws IOException {
                if (operands.get(0) instanceof COSString cosString) {
                    cosString.setValue(cleanedText.getBytes());
                } else if (operands.get(0) instanceof COSArray cosArray) {
                    cosArray.clear();
                    cosArray.add(new COSString(cleanedText));
                }
            }
        }.processPage(page);
    }
}

使用方式

try (PDDocument document = PDDocument.load(new File("your-document.pdf"))) {
    for (PDPage page : document.getPages()) {
        PlaceholderRemover.removePlaceholders(document, page);
    }
    document.save("cleaned-document.pdf");
} catch (IOException e) {
    e.printStackTrace();
}

核心思路

  1. 跟踪文本上下文:用缓冲区累积连续文本操作内容,避免占位符被拆分导致漏处理。
  2. 识别占位符状态:判断文本是否处于占位符的起始、中间或结束状态,决定是否跳过写入。
  3. 保留布局一致性:将占位符替换为同等长度的空格,避免移除文本后页面布局错乱。
  4. 兼容多操作类型:覆盖所有PDF文本绘制操作,适配不同生成工具的输出格式。

内容的提问来源于stack exchange,提问作者Vas K

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 04:27:33