如何用PDFBox移除PDF指定坐标区域的文本?
Hey there, I get your frustration—covering text with a black rectangle works for visual purposes, but it doesn't truly remove the underlying content which poses a security risk. Since PDFBox dropped text replacement support and you can't use iText, here are two practical, reliable workarounds:
1. Regenerate the PDF with Filtered Text
This approach involves extracting all content from the original PDF, excluding the text in your target region, then rebuilding a new PDF from scratch. It's straightforward and less error-prone for most cases.
Step-by-Step Implementation
First, create a custom PDFTextStripper subclass to filter out text in your target area:
import org.apache.pdfbox.pdmodel.PDDocument; import org.apache.pdfbox.text.PDFTextStripper; import org.apache.pdfbox.text.TextPosition; import java.io.IOException; import java.util.List; public class RegionFilteredTextStripper extends PDFTextStripper { private final float targetX1, targetX2, targetY1, targetY2; private final int targetPageNumber; public RegionFilteredTextStripper(float x1, float x2, float y1, float y2, int pageNum) throws IOException { this.targetX1 = x1; this.targetX2 = x2; this.targetY1 = y1; this.targetY2 = y2; this.targetPageNumber = pageNum; } @Override protected void writeString(String text, List<TextPosition> textPositions) throws IOException { // Skip filtering if we're not on the target page if (getCurrentPageNo() != targetPageNumber) { super.writeString(text, textPositions); return; } // Check if any part of the text overlaps with the target region boolean isOutsideRegion = true; for (TextPosition tp : textPositions) { float textX = tp.getX(); float textY = tp.getY(); float textEndX = textX + tp.getWidth(); float textEndY = textY + tp.getHeight(); // Check for overlap (inverse of "no overlap") if (!(textEndX < targetX1 || textX > targetX2 || textEndY < targetY1 || textY > targetY2)) { isOutsideRegion = false; break; } } // Only write text that's outside the target region if (isOutsideRegion) { super.writeString(text, textPositions); } } }
Next, use this stripper to extract filtered text, then combine it with non-text elements (images, graphics) from the original PDF to build a new document:
import org.apache.pdfbox.pdmodel.PDDocument; import org.apache.pdfbox.pdmodel.PDPage; import org.apache.pdfbox.pdmodel.PDPageContentStream; import org.apache.pdfbox.pdmodel.font.PDType1Font; import java.io.File; import java.io.IOException; public class PdfTextRemover { public static void main(String[] args) throws IOException { // Define your target region and page float x1 = 100, x2 = 300, y1 = 200, y2 = 250; int targetPage = 1; // PDFBox pages are 1-indexed in the stripper // Load original document try (PDDocument originalDoc = PDDocument.load(new File("input.pdf")); PDDocument newDoc = new PDDocument()) { // Copy all pages first for (PDPage page : originalDoc.getPages()) { newDoc.addPage(page); } // Extract filtered text from target page RegionFilteredTextStripper stripper = new RegionFilteredTextStripper(x1, x2, y1, y2, targetPage); stripper.setStartPage(targetPage); stripper.setEndPage(targetPage); String filteredText = stripper.getText(originalDoc); // Overwrite the target page's content with filtered text + original non-text elements PDPage targetPageObj = newDoc.getPage(targetPage - 1); // PDFBox pages are 0-indexed here try (PDPageContentStream contentStream = new PDPageContentStream(newDoc, targetPageObj, PDPageContentStream.AppendMode.OVERWRITE, true, true)) { // Add filtered text (adjust coordinates as needed based on your PDF's origin) contentStream.beginText(); contentStream.setFont(PDType1Font.HELVETICA, 12); contentStream.newLineAtOffset(50, 700); // Example position, adjust to match original layout contentStream.showText(filteredText); contentStream.endText(); // Optional: Add the black rectangle for extra security contentStream.setNonStrokingColor(java.awt.Color.BLACK); contentStream.addRect(x1, y1, x2 - x1, y2 - y1); contentStream.fill(); } // Save the modified document newDoc.save("output.pdf"); } } }
2. Directly Modify the Page's Content Stream
For more control and better performance (especially for large PDFs), you can parse the PDF's content stream to remove text-drawing commands that fall within your target region. This is more low-level but removes the text entirely.
Key Implementation Notes
PDF content streams use operators like Tj (single text string) and TJ (text array) to draw text. You'll need to:
- Track the current text matrix (to calculate text positions)
- Track the current font (to calculate text width)
- Filter out any text-drawing commands that overlap with your target region
Here's a simplified example of parsing the content stream:
import org.apache.pdfbox.contentstream.PDFStreamParser; import org.apache.pdfbox.contentstream.operator.Operator; import org.apache.pdfbox.cos.COSArray; import org.apache.pdfbox.cos.COSString; import org.apache.pdfbox.pdmodel.PDDocument; import org.apache.pdfbox.pdmodel.PDPage; import org.apache.pdfbox.pdmodel.PDStream; import org.apache.pdfbox.util.Matrix; import java.io.File; import java.io.IOException; import java.util.ArrayList; import java.util.List; public class ContentStreamModifier { public static void main(String[] args) throws IOException { float x1 = 100, x2 = 300, y1 = 200, y2 = 250; int targetPageIndex = 0; // 0-indexed try (PDDocument document = PDDocument.load(new File("input.pdf"))) { PDPage page = document.getPage(targetPageIndex); PDStream stream = page.getContents(); PDFStreamParser parser = new PDFStreamParser(stream); parser.parse(); List<Object> tokens = parser.getTokens(); List<Object> newTokens = new ArrayList<>(); Matrix currentTextMatrix = Matrix.IDENTITY; for (int i = 0; i < tokens.size(); i++) { Object token = tokens.get(i); if (token instanceof Operator) { Operator op = (Operator) token; switch (op.getName()) { case "Tm": // Update text matrix COSArray matrixArray = (COSArray) tokens.get(i - 3); currentTextMatrix = Matrix.createMatrix(matrixArray); // Keep the matrix command newTokens.add(tokens.get(i-3)); newTokens.add(tokens.get(i-2)); newTokens.add(tokens.get(i-1)); newTokens.add(op); i += 3; break; case "Tj": // Draw single text string COSString textString = (COSString) tokens.get(i - 1); float textX = currentTextMatrix.getTranslateX(); float textY = currentTextMatrix.getTranslateY(); // Calculate text width (you'll need to get current font for accurate measurement) float textWidth = textString.getString().length() * 6; // Simplified, adjust based on font // Check if text is outside target region if (textX + textWidth < x1 || textX > x2 || textY + 12 < y1 || textY > y2) { newTokens.add(textString); newTokens.add(op); } i += 1; break; // Handle other text operators like TJ, Tf, etc. as needed default: newTokens.add(token); break; } } else { newTokens.add(token); } } // Replace the page's content stream with filtered tokens PDStream newStream = new PDStream(document); try (var os = newStream.createOutputStream()) { var writer = new org.apache.pdfbox.contentstream.ContentStreamWriter(os); writer.writeTokens(newTokens); } page.setContents(newStream); // Optional: Add black rectangle try (PDPageContentStream cs = new PDPageContentStream(document, page, PDPageContentStream.AppendMode.APPEND, true)) { cs.setNonStrokingColor(java.awt.Color.BLACK); cs.addRect(x1, y1, x2 - x1, y2 - y1); cs.fill(); } document.save("output.pdf"); } } }
Important Notes
- Backup First: Always work on a copy of your original PDF to avoid data loss.
- Complex PDFs: For PDFs with nested content streams, rotated pages, or custom fonts, you'll need to expand the logic to handle those cases (e.g., tracking font sizes, handling rotation matrices).
- Testing: Validate the output PDF to ensure no unintended content was removed and the layout remains intact.
内容的提问来源于stack exchange,提问作者cagy

