如何在Java中使用PDFBox或其他库压缩PDF文件大小?
Java项目中压缩PDF文件体积的实用方案
针对你使用iTextPDF/PDFBox处理PDF时的体积优化需求,以下是具体的实现方案和代码示例:
一、iTextPDF 优化方案
1. 基础全压缩(优化你的现有代码)
你当前的代码已经启用了setFullCompression(),这会开启对象流和交叉引用流的压缩,是iText的基础全压缩机制。可以补充设置压缩级别(0-9,9为最高压缩比)进一步优化:
PdfReader reader = new PdfReader(inputPdf); ByteArrayOutputStream outputStream = new ByteArrayOutputStream(); try { PdfStamper stamper = new PdfStamper(reader, outputStream); // 启用全压缩 stamper.setFullCompression(); // 设置最高压缩级别 stamper.getWriter().setCompressionLevel(9); stamper.close(); } finally { reader.close(); } return outputStream.toByteArray();
2. 压缩PDF中的图片(最有效降体积手段)
PDF体积过大往往源于未优化的图片,你可以遍历PDF中的图片资源,通过降低分辨率、转换格式(如PNG转JPEG)、调整质量来压缩:
PdfReader reader = new PdfReader(inputPdf); ByteArrayOutputStream outputStream = new ByteArrayOutputStream(); PdfStamper stamper = new PdfStamper(reader, outputStream); // 遍历所有页面 for (int i = 1; i <= reader.getNumberOfPages(); i++) { PdfDictionary pageDict = reader.getPageN(i); PdfDictionary resources = pageDict.getAsDict(PdfName.RESOURCES); if (resources == null) continue; // 遍历图片资源 PdfDictionary xObjects = resources.getAsDict(PdfName.XOBJECT); if (xObjects == null) continue; for (PdfName key : xObjects.getKeys()) { PdfObject obj = xObjects.get(key); if (obj instanceof PdfIndirectReference) { PdfDictionary imgDict = (PdfDictionary) PdfReader.getPdfObject(obj); PdfName subtype = imgDict.getAsName(PdfName.SUBTYPE); if (PdfName.IMAGE.equals(subtype)) { // 获取图片并压缩(示例:将PNG转为质量70%的JPEG) PdfImageObject image = new PdfImageObject(imgDict); BufferedImage bufferedImage = image.getBufferedImage(); ByteArrayOutputStream imgBaos = new ByteArrayOutputStream(); ImageIO.write(bufferedImage, "JPEG", new javax.imageio.stream.MemoryCacheImageOutputStream(imgBaos)); // 替换原图片 PdfStream compressedStream = new PdfStream(imgBaos.toByteArray()); compressedStream.putAll(imgDict); compressedStream.put(PdfName.FILTER, PdfName.DCTDECODE); compressedStream.put(PdfName.BITSPERCOMPONENT, new PdfNumber(8)); compressedStream.put(PdfName.COLORSPACE, PdfName.DEVICERGB); xObjects.put(key, stamper.getWriter().addToBody(compressedStream).getIndirectReference()); } } } } stamper.setFullCompression(); stamper.getWriter().setCompressionLevel(9); stamper.close(); reader.close(); return outputStream.toByteArray();
3. 移除冗余内容
移除PDF中不需要的元素(如注释、未使用字体、元数据)也能减少体积:
// 在stamper创建后添加以下代码 // 移除所有注释 for (int i = 1; i <= reader.getNumberOfPages(); i++) { PdfDictionary pageDict = reader.getPageN(i); pageDict.remove(PdfName.ANNOTS); } // 移除未使用的字体 reader.removeUnusedFonts();
二、PDFBox 优化方案
如果你切换使用PDFBox,也可以通过以下方式压缩:
1. 基础文档压缩
启用内容流和交叉引用流的压缩:
try (PDDocument document = PDDocument.load(new File(inputPdfPath))) { SaveOptions saveOptions = new SaveOptions(); // 压缩内容流 saveOptions.setCompressContentStreams(true); // 压缩交叉引用流 saveOptions.setCompressXrefStreams(true); // 设置最高压缩级别 saveOptions.setCompressionLevel(9); document.save(outputPdfPath, saveOptions); } catch (IOException e) { e.printStackTrace(); }
2. 图片压缩与优化
遍历并压缩文档中的图片:
try (PDDocument document = PDDocument.load(new File(inputPdfPath))) { for (PDPage page : document.getPages()) { PDResources resources = page.getResources(); for (COSName name : resources.getXObjectNames()) { PDXObject xobject = resources.getXObject(name); if (xobject instanceof PDImageXObject) { PDImageXObject image = (PDImageXObject) xobject; // 降低分辨率至150DPI(适合打印/屏幕显示) image.setDPI(150, 150); // 将PNG转为JPEG压缩(质量70%) if ("png".equalsIgnoreCase(image.getSuffix())) { BufferedImage bufferedImage = image.getImage(); ByteArrayOutputStream baos = new ByteArrayOutputStream(); ImageWriteParam param = ImageIO.getImageWritersByFormatName("JPEG").next().getDefaultWriteParam(); param.setCompressionMode(ImageWriteParam.MODE_EXPLICIT); param.setCompressionQuality(0.7f); ImageIO.write(bufferedImage, "JPEG", new javax.imageio.stream.MemoryCacheImageOutputStream(baos)); PDImageXObject compressedImage = PDImageXObject.createFromByteArray(document, baos.toByteArray(), "JPEG"); resources.putXObject(name, compressedImage); } } } } // 清理未使用的资源 document.cleanUp(); SaveOptions saveOptions = new SaveOptions(); saveOptions.setCompressContentStreams(true); saveOptions.setCompressionLevel(9); document.save(outputPdfPath, saveOptions); } catch (IOException e) { e.printStackTrace(); }
3. 移除冗余元素
try (PDDocument document = PDDocument.load(new File(inputPdfPath))) { // 移除表单域 document.getDocumentCatalog().setAcroForm(null); // 移除所有页面注释 for (PDPage page : document.getPages()) { page.setAnnotations(new ArrayList<>()); } // 清除元数据(如作者、标题) PDDocumentInformation info = document.getDocumentInformation(); info.setAuthor(null); info.setTitle(null); info.setSubject(null); document.cleanUp(); document.save(outputPdfPath); } catch (IOException e) { e.printStackTrace(); }
内容的提问来源于stack exchange,提问作者Kishor
相关产品推荐
相关产品推荐

