使用PDFBox 2.x移除PDF二维码后Adobe Reader报错及文件变大求助
问题分析与解决方案
问题概述
使用PDFBox 2.x移除PDF中的二维码后,Adobe Reader打开修改后的文档时会弹出「此页面存在错误,Acrobat可能无法正确显示页面」的警告,且文件体积从6.8MB增至8.1MB,预期无警告弹窗且文件大小不增加。
核心问题排查
从提供的代码中,可定位出两个关键问题:
1. 资源字典操作不规范
直接操作底层COS对象移除图片资源:
((COSDictionary) pdResources.getCOSObject().getDictionaryObject(COSName.XOBJECT)) .removeItem(propertyName);
这种方式绕过了PDFBox的API封装,会导致资源字典内部状态不一致,PDFBox无法正确跟踪资源变化,触发Acrobat的文档验证错误。
2. 内容流处理存在漏洞
处理Do操作符时仅移除前一个COSName token,若该指令前后存在状态设置(如矩阵变换、颜色配置),会导致内容流语法/逻辑不完整;同时用新流直接替换原页面所有内容流,可能破坏原有文档结构。
3. 文件体积增大原因
- 导入页面时未启用压缩,新生成的内容流无压缩处理;
- 原文档冗余资源未被清理,新文档保留了未使用的对象。
修复方案
1. 规范移除资源项
使用PDFBox官方APIPDResources.removeXObject()替代直接操作COS对象,确保资源字典状态一致:
pdResources.removeXObject(propertyName);
2. 正确处理内容流
完整跳过Do操作符及其关联的对象名,同时保留原有内容流结构(避免替换整个流):
PDFStreamParser parser = new PDFStreamParser(page); parser.parse(); List<Object> tokens = parser.getTokens(); List<Object> newTokens = new ArrayList<>(); int i = 0; while (i < tokens.size()) { Object token = tokens.get(i); if (token instanceof Operator) { Operator op = (Operator) token; if ("Do".equals(op.getName()) && i > 0) { COSName cosName = (COSName) tokens.get(i - 1); if (cosName.getName().equals(qrCodeCosName)) { // 跳过对象名与Do操作符 i += 2; continue; } } } newTokens.add(token); i++; } // 保留原有内容流结构,采用追加模式 PDStream existingStream = page.getContents(); if (existingStream != null) { try (OutputStream out = existingStream.createOutputStreamAppend()) { ContentStreamWriter writer = new ContentStreamWriter(out); writer.writeTokens(newTokens); } } else { PDStream newContents = new PDStream(document); try (OutputStream out = newContents.createOutputStream()) { ContentStreamWriter writer = new ContentStreamWriter(out); writer.writeTokens(newTokens); } page.setContents(newContents); }
3. 启用压缩并清理冗余对象
保存文档前清理未使用对象,同时启用压缩:
// 清理冗余对象 newDocument.cleanUp(); // 启用压缩保存 newDocument.save(aBarcodeVO.getDestinationFilePath(true), SaveOptions.builder().setCompress(true).build());
完整修复后核心代码
pdDocument = PDDocument.load(new File(aBarcodeVO.getSourceFilePath())); newDocument = new PDDocument(); for (int pageCount = 0; pageCount < pdDocument.getNumberOfPages(); pageCount++) { PDPage pdPage = newDocument.importPage(pdDocument.getPage(pageCount)); String imgUniqueId = aBarcodeVO.getImgUniqueId().concat(String.valueOf(pageCount)); boolean hasQRCodeOnPage = removeQRCodeImage(newDocument, pdPage, imgUniqueId); qRCodePageList.add(hasQRCodeOnPage); } if(qRCodePageList.contains(true)) { newDocument.cleanUp(); newDocument.save(aBarcodeVO.getDestinationFilePath(true), SaveOptions.builder().setCompress(true).build()); } newDocument.close(); pdDocument.close(); // 修复后的removeQRCodeImage方法 public static boolean removeQRCodeImage(PDDocument document, PDPage page, String imgUniqueId) throws Exception { String qrCodeCosName = null; PDResources pdResources = page.getResources(); boolean hasQRCodeOnPage=false; Iterator<COSName> xObjectNames = pdResources.getXObjectNames().iterator(); while (xObjectNames.hasNext()) { COSName propertyName = xObjectNames.next(); if (!pdResources.isImageXObject(propertyName)) { continue; } PDXObject o; try { o = pdResources.getXObject(propertyName); if (o instanceof PDImageXObject) { PDImageXObject pdImageXObject = (PDImageXObject) o; if (pdImageXObject.getMetadata() != null) { DomXmpParser xmpParser = new DomXmpParser(); XMPMetadata xmpMetadata = xmpParser.parse(pdImageXObject.getMetadata().toByteArray()); if(xmpMetadata.getDublinCoreSchema()!=null && StringUtils.isNoneBlank(xmpMetadata.getDublinCoreSchema().getTitle())&&xmpMetadata.getDublinCoreSchema().getTitle().contains("_barcodeimg_")) { pdResources.removeXObject(propertyName); log.debug("propertyName REMOVED--"+propertyName.getName()); qrCodeCosName = propertyName.getName(); hasQRCodeOnPage=true; } } } } catch (IOException e) { log.error("Exception in removeQRCodeImage() while extracting QR image:" + e, e); } } if (qrCodeCosName == null) { return hasQRCodeOnPage; } PDFStreamParser parser = new PDFStreamParser(page); parser.parse(); List<Object> tokens = parser.getTokens(); log.debug("original tokens size" + tokens.size()); List<Object> newTokens = new ArrayList<>(); int i = 0; while (i < tokens.size()) { Object token = tokens.get(i); if (token instanceof Operator) { Operator op = (Operator) token; if ("Do".equals(op.getName()) && i > 0) { COSName cosName = (COSName) tokens.get(i - 1); if (cosName.getName().equals(qrCodeCosName)) { i += 2; continue; } } } newTokens.add(token); i++; } log.debug("tokens size" + newTokens.size()); PDStream existingStream = page.getContents(); if (existingStream != null) { try (OutputStream out = existingStream.createOutputStreamAppend()) { ContentStreamWriter writer = new ContentStreamWriter(out); writer.writeTokens(newTokens); } } else { PDStream newContents = new PDStream(document); try (OutputStream out = newContents.createOutputStream()) { ContentStreamWriter writer = new ContentStreamWriter(out); writer.writeTokens(newTokens); } page.setContents(newContents); } return hasQRCodeOnPage; }
验证结果
- Adobe Reader打开修改后的文档不再弹出错误警告;
- 文件体积因压缩和冗余资源清理,会减小或接近原文件大小。
内容的提问来源于stack exchange,提问作者Ayush
相关产品推荐
相关产品推荐

