使用iText 7.1移除PDF链接失败求助:三种尝试均无效
Hey there! Let's tackle this problem where your iText 7.1 code isn't removing links from PDFs as expected. I've looked over your three attempts, and I'll break down why they might not be working, plus give you a solid solution that covers all common link types in PDFs.
Why Your Existing Approaches Might Be Failing
Let's quickly go through your three code samples to spot potential gaps:
- Sample 1 (Remove by Annotation Type): This only targets
PdfLinkAnnotationinstances, but misses links embedded in AcroForm fields (like button fields with URI actions) or links stored as indirect annotations (sincegetAnnotations()defaults to skipping indirect references). - Sample 2 (Remove by Subtype): You're filtering only
URIandGoToRactions, but there are other link types (like localGoTojumps) you're not catching. Plus, it has the same AcroForm/indirect annotation blind spots as Sample 1. - Sample 3 (Remove All Annotations): This only removes page-level annotations stored in the
Annotsdictionary. It won't touch links in AcroForm fields or links embedded directly in the page's content stream (via/Aoperators). Also, you didn't callpdfPage.flush()to save the removal change to the document.
Solution: Cover All Link Types
To fully remove all links from a PDF, you need to handle three scenarios: page link annotations, AcroForm link fields, and (rarely) content-stream embedded links. Here's a comprehensive implementation:
Step 1: Remove Link Annotations and AcroForm Link Fields
This code handles the two most common link types:
import com.itextpdf.kernel.pdf.*; import com.itextpdf.forms.PdfAcroForm; import com.itextpdf.forms.fields.PdfFormField; import java.util.ArrayList; import java.util.List; import java.util.Map; public class PdfLinkRemover { public static void main(String[] args) { String src = "test-with-links.pdf"; String dest = "test-no-links.pdf"; try (PdfReader reader = new PdfReader(src); PdfWriter writer = new PdfWriter(dest); PdfDocument pdfDoc = new PdfDocument(reader, writer)) { // Handle AcroForm fields with link actions PdfAcroForm form = PdfAcroForm.getAcroForm(pdfDoc, false); if (form != null) { Map<String, PdfFormField> fields = form.getFormFields(); List<PdfFormField> fieldsToRemove = new ArrayList<>(); for (PdfFormField field : fields.values()) { PdfAction action = field.getAction(); if (action != null) { PdfName actionType = action.getSubtype(); // Check for all common link action types if (PdfName.URI.equals(actionType) || PdfName.GoToR.equals(actionType) || PdfName.GoTo.equals(actionType)) { // Option 1: Clear the action (keep the field) field.setAction(null); // Option 2: Remove the field entirely (uncomment if needed) // fieldsToRemove.add(field); } } } // Remove marked fields if you chose Option 2 for (PdfFormField field : fieldsToRemove) { form.removeField(field.getFieldName()); } } // Handle page link annotations (including indirect annotations) for (int pageNum = 1; pageNum <= pdfDoc.getNumberOfPages(); pageNum++) { PdfPage page = pdfDoc.getPage(pageNum); // Get ALL annotations (pass true to include indirect references) List<PdfAnnotation> annotations = page.getAnnotations(true); if (annotations != null && !annotations.isEmpty()) { List<PdfAnnotation> linksToRemove = new ArrayList<>(); for (PdfAnnotation annot : annotations) { // Check both type and subtype to cover all link annotations if (annot instanceof PdfLinkAnnotation || PdfName.Link.equals(annot.getSubtype())) { linksToRemove.add(annot); } } // Remove collected links safely (avoids ConcurrentModificationException) for (PdfAnnotation linkAnnot : linksToRemove) { page.removeAnnotation(linkAnnot); } } // Ensure page changes are saved page.flush(); } } catch (Exception e) { e.printStackTrace(); } } }
Step 2: Handle Content-Stream Embedded Links (Rare Case)
If your PDF has links embedded directly in the page content stream (not as annotations or form fields), you'll need to parse and rewrite the content stream to remove action references. This is more complex, but here's a basic implementation using PdfCanvasProcessor:
import com.itextpdf.kernel.pdf.*; import com.itextpdf.kernel.pdf.canvas.PdfCanvasProcessor; import com.itextpdf.kernel.pdf.canvas.parser.IEventListener; import com.itextpdf.kernel.pdf.canvas.parser.data.IEventData; import com.itextpdf.kernel.pdf.canvas.parser.EventType; import com.itextpdf.kernel.pdf.canvas.parser.listener.RenderListener; import com.itextpdf.kernel.pdf.canvas.parser.data.TextRenderInfo; import com.itextpdf.kernel.pdf.canvas.parser.data.ImageRenderInfo; import com.itextpdf.kernel.pdf.canvas.parser.data.PathRenderInfo; import java.util.Set; import java.util.HashSet; public class PdfContentStreamLinkRemover { public static void main(String[] args) { String src = "test-with-links.pdf"; String dest = "test-no-links.pdf"; try (PdfReader reader = new PdfReader(src); PdfWriter writer = new PdfWriter(dest); PdfDocument pdfDoc = new PdfDocument(reader, writer)) { for (int pageNum = 1; pageNum <= pdfDoc.getNumberOfPages(); pageNum++) { PdfPage page = pdfDoc.getPage(pageNum); // Create a new content stream to write cleaned content PdfCanvas newCanvas = new PdfCanvas(page.newContentStreamBefore(), page.getResources(), pdfDoc); // Process the original content stream, skipping actions PdfCanvasProcessor processor = new PdfCanvasProcessor(new RenderListener() { @Override public void renderText(TextRenderInfo renderInfo) { renderInfo.renderText(newCanvas); } @Override public void renderImage(ImageRenderInfo renderInfo) { renderInfo.renderImage(newCanvas); } @Override public void renderPath(PathRenderInfo renderInfo) { renderInfo.renderPath(newCanvas); } @Override public void clipPath(int rule) { // No-op, or replicate if needed } }) { @Override public void eventOccurred(IEventData data, EventType type) { // Skip action events entirely if (!EventType.ACTION.equals(type)) { super.eventOccurred(data, type); } } @Override public Set<EventType> getSupportedEvents() { Set<EventType> events = new HashSet<>(super.getSupportedEvents()); events.add(EventType.ACTION); return events; } }; processor.processPageContent(page); // Replace the original content stream with the cleaned one page.removePageStream(); page.flush(); } } catch (Exception e) { e.printStackTrace(); } } }
Key Notes
- Always use try-with-resources for
PdfReader,PdfWriter, andPdfDocumentto ensure resources are properly closed. - When removing elements from a list (like annotations), collect them first then remove to avoid
ConcurrentModificationException. - Call
page.flush()after modifying page content to ensure changes are persisted to the document.
内容的提问来源于stack exchange,提问作者Jack D.

