You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用PDFBox移除PDF指定坐标区域的文本?

Solutions to Remove Underlying Text in PDF with PDFBox (No iText Allowed)

Hey there, I get your frustration—covering text with a black rectangle works for visual purposes, but it doesn't truly remove the underlying content which poses a security risk. Since PDFBox dropped text replacement support and you can't use iText, here are two practical, reliable workarounds:

1. Regenerate the PDF with Filtered Text

This approach involves extracting all content from the original PDF, excluding the text in your target region, then rebuilding a new PDF from scratch. It's straightforward and less error-prone for most cases.

Step-by-Step Implementation

First, create a custom PDFTextStripper subclass to filter out text in your target area:

import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.text.PDFTextStripper;
import org.apache.pdfbox.text.TextPosition;
import java.io.IOException;
import java.util.List;

public class RegionFilteredTextStripper extends PDFTextStripper {
    private final float targetX1, targetX2, targetY1, targetY2;
    private final int targetPageNumber;

    public RegionFilteredTextStripper(float x1, float x2, float y1, float y2, int pageNum) throws IOException {
        this.targetX1 = x1;
        this.targetX2 = x2;
        this.targetY1 = y1;
        this.targetY2 = y2;
        this.targetPageNumber = pageNum;
    }

    @Override
    protected void writeString(String text, List<TextPosition> textPositions) throws IOException {
        // Skip filtering if we're not on the target page
        if (getCurrentPageNo() != targetPageNumber) {
            super.writeString(text, textPositions);
            return;
        }

        // Check if any part of the text overlaps with the target region
        boolean isOutsideRegion = true;
        for (TextPosition tp : textPositions) {
            float textX = tp.getX();
            float textY = tp.getY();
            float textEndX = textX + tp.getWidth();
            float textEndY = textY + tp.getHeight();

            // Check for overlap (inverse of "no overlap")
            if (!(textEndX < targetX1 || textX > targetX2 || textEndY < targetY1 || textY > targetY2)) {
                isOutsideRegion = false;
                break;
            }
        }

        // Only write text that's outside the target region
        if (isOutsideRegion) {
            super.writeString(text, textPositions);
        }
    }
}

Next, use this stripper to extract filtered text, then combine it with non-text elements (images, graphics) from the original PDF to build a new document:

import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.pdmodel.PDPage;
import org.apache.pdfbox.pdmodel.PDPageContentStream;
import org.apache.pdfbox.pdmodel.font.PDType1Font;
import java.io.File;
import java.io.IOException;

public class PdfTextRemover {
    public static void main(String[] args) throws IOException {
        // Define your target region and page
        float x1 = 100, x2 = 300, y1 = 200, y2 = 250;
        int targetPage = 1; // PDFBox pages are 1-indexed in the stripper

        // Load original document
        try (PDDocument originalDoc = PDDocument.load(new File("input.pdf"));
             PDDocument newDoc = new PDDocument()) {

            // Copy all pages first
            for (PDPage page : originalDoc.getPages()) {
                newDoc.addPage(page);
            }

            // Extract filtered text from target page
            RegionFilteredTextStripper stripper = new RegionFilteredTextStripper(x1, x2, y1, y2, targetPage);
            stripper.setStartPage(targetPage);
            stripper.setEndPage(targetPage);
            String filteredText = stripper.getText(originalDoc);

            // Overwrite the target page's content with filtered text + original non-text elements
            PDPage targetPageObj = newDoc.getPage(targetPage - 1); // PDFBox pages are 0-indexed here
            try (PDPageContentStream contentStream = new PDPageContentStream(newDoc, targetPageObj, PDPageContentStream.AppendMode.OVERWRITE, true, true)) {
                // Add filtered text (adjust coordinates as needed based on your PDF's origin)
                contentStream.beginText();
                contentStream.setFont(PDType1Font.HELVETICA, 12);
                contentStream.newLineAtOffset(50, 700); // Example position, adjust to match original layout
                contentStream.showText(filteredText);
                contentStream.endText();

                // Optional: Add the black rectangle for extra security
                contentStream.setNonStrokingColor(java.awt.Color.BLACK);
                contentStream.addRect(x1, y1, x2 - x1, y2 - y1);
                contentStream.fill();
            }

            // Save the modified document
            newDoc.save("output.pdf");
        }
    }
}

2. Directly Modify the Page's Content Stream

For more control and better performance (especially for large PDFs), you can parse the PDF's content stream to remove text-drawing commands that fall within your target region. This is more low-level but removes the text entirely.

Key Implementation Notes

PDF content streams use operators like Tj (single text string) and TJ (text array) to draw text. You'll need to:

  • Track the current text matrix (to calculate text positions)
  • Track the current font (to calculate text width)
  • Filter out any text-drawing commands that overlap with your target region

Here's a simplified example of parsing the content stream:

import org.apache.pdfbox.contentstream.PDFStreamParser;
import org.apache.pdfbox.contentstream.operator.Operator;
import org.apache.pdfbox.cos.COSArray;
import org.apache.pdfbox.cos.COSString;
import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.pdmodel.PDPage;
import org.apache.pdfbox.pdmodel.PDStream;
import org.apache.pdfbox.util.Matrix;
import java.io.File;
import java.io.IOException;
import java.util.ArrayList;
import java.util.List;

public class ContentStreamModifier {
    public static void main(String[] args) throws IOException {
        float x1 = 100, x2 = 300, y1 = 200, y2 = 250;
        int targetPageIndex = 0; // 0-indexed

        try (PDDocument document = PDDocument.load(new File("input.pdf"))) {
            PDPage page = document.getPage(targetPageIndex);
            PDStream stream = page.getContents();
            PDFStreamParser parser = new PDFStreamParser(stream);
            parser.parse();
            List<Object> tokens = parser.getTokens();
            List<Object> newTokens = new ArrayList<>();

            Matrix currentTextMatrix = Matrix.IDENTITY;

            for (int i = 0; i < tokens.size(); i++) {
                Object token = tokens.get(i);
                if (token instanceof Operator) {
                    Operator op = (Operator) token;
                    switch (op.getName()) {
                        case "Tm": // Update text matrix
                            COSArray matrixArray = (COSArray) tokens.get(i - 3);
                            currentTextMatrix = Matrix.createMatrix(matrixArray);
                            // Keep the matrix command
                            newTokens.add(tokens.get(i-3));
                            newTokens.add(tokens.get(i-2));
                            newTokens.add(tokens.get(i-1));
                            newTokens.add(op);
                            i += 3;
                            break;
                        case "Tj": // Draw single text string
                            COSString textString = (COSString) tokens.get(i - 1);
                            float textX = currentTextMatrix.getTranslateX();
                            float textY = currentTextMatrix.getTranslateY();
                            // Calculate text width (you'll need to get current font for accurate measurement)
                            float textWidth = textString.getString().length() * 6; // Simplified, adjust based on font

                            // Check if text is outside target region
                            if (textX + textWidth < x1 || textX > x2 || textY + 12 < y1 || textY > y2) {
                                newTokens.add(textString);
                                newTokens.add(op);
                            }
                            i += 1;
                            break;
                        // Handle other text operators like TJ, Tf, etc. as needed
                        default:
                            newTokens.add(token);
                            break;
                    }
                } else {
                    newTokens.add(token);
                }
            }

            // Replace the page's content stream with filtered tokens
            PDStream newStream = new PDStream(document);
            try (var os = newStream.createOutputStream()) {
                var writer = new org.apache.pdfbox.contentstream.ContentStreamWriter(os);
                writer.writeTokens(newTokens);
            }
            page.setContents(newStream);

            // Optional: Add black rectangle
            try (PDPageContentStream cs = new PDPageContentStream(document, page, PDPageContentStream.AppendMode.APPEND, true)) {
                cs.setNonStrokingColor(java.awt.Color.BLACK);
                cs.addRect(x1, y1, x2 - x1, y2 - y1);
                cs.fill();
            }

            document.save("output.pdf");
        }
    }
}

Important Notes

  • Backup First: Always work on a copy of your original PDF to avoid data loss.
  • Complex PDFs: For PDFs with nested content streams, rotated pages, or custom fonts, you'll need to expand the logic to handle those cases (e.g., tracking font sizes, handling rotation matrices).
  • Testing: Validate the output PDF to ensure no unintended content was removed and the layout remains intact.

内容的提问来源于stack exchange,提问作者cagy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:38:21