You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Java中使用PDFBox或其他库压缩PDF文件大小?

Java项目中压缩PDF文件体积的实用方案

针对你使用iTextPDF/PDFBox处理PDF时的体积优化需求,以下是具体的实现方案和代码示例:


一、iTextPDF 优化方案

1. 基础全压缩(优化你的现有代码)

你当前的代码已经启用了setFullCompression(),这会开启对象流和交叉引用流的压缩,是iText的基础全压缩机制。可以补充设置压缩级别(0-9,9为最高压缩比)进一步优化:

PdfReader reader = new PdfReader(inputPdf);
ByteArrayOutputStream outputStream = new ByteArrayOutputStream();
try {
    PdfStamper stamper = new PdfStamper(reader, outputStream);
    // 启用全压缩
    stamper.setFullCompression();
    // 设置最高压缩级别
    stamper.getWriter().setCompressionLevel(9);
    stamper.close();
} finally {
    reader.close();
}
return outputStream.toByteArray();

2. 压缩PDF中的图片(最有效降体积手段)

PDF体积过大往往源于未优化的图片,你可以遍历PDF中的图片资源,通过降低分辨率、转换格式(如PNG转JPEG)、调整质量来压缩:

PdfReader reader = new PdfReader(inputPdf);
ByteArrayOutputStream outputStream = new ByteArrayOutputStream();
PdfStamper stamper = new PdfStamper(reader, outputStream);

// 遍历所有页面
for (int i = 1; i <= reader.getNumberOfPages(); i++) {
    PdfDictionary pageDict = reader.getPageN(i);
    PdfDictionary resources = pageDict.getAsDict(PdfName.RESOURCES);
    if (resources == null) continue;
    
    // 遍历图片资源
    PdfDictionary xObjects = resources.getAsDict(PdfName.XOBJECT);
    if (xObjects == null) continue;
    
    for (PdfName key : xObjects.getKeys()) {
        PdfObject obj = xObjects.get(key);
        if (obj instanceof PdfIndirectReference) {
            PdfDictionary imgDict = (PdfDictionary) PdfReader.getPdfObject(obj);
            PdfName subtype = imgDict.getAsName(PdfName.SUBTYPE);
            
            if (PdfName.IMAGE.equals(subtype)) {
                // 获取图片并压缩(示例:将PNG转为质量70%的JPEG)
                PdfImageObject image = new PdfImageObject(imgDict);
                BufferedImage bufferedImage = image.getBufferedImage();
                
                ByteArrayOutputStream imgBaos = new ByteArrayOutputStream();
                ImageIO.write(bufferedImage, "JPEG", new javax.imageio.stream.MemoryCacheImageOutputStream(imgBaos));
                
                // 替换原图片
                PdfStream compressedStream = new PdfStream(imgBaos.toByteArray());
                compressedStream.putAll(imgDict);
                compressedStream.put(PdfName.FILTER, PdfName.DCTDECODE);
                compressedStream.put(PdfName.BITSPERCOMPONENT, new PdfNumber(8));
                compressedStream.put(PdfName.COLORSPACE, PdfName.DEVICERGB);
                
                xObjects.put(key, stamper.getWriter().addToBody(compressedStream).getIndirectReference());
            }
        }
    }
}

stamper.setFullCompression();
stamper.getWriter().setCompressionLevel(9);
stamper.close();
reader.close();
return outputStream.toByteArray();

3. 移除冗余内容

移除PDF中不需要的元素(如注释、未使用字体、元数据)也能减少体积:

// 在stamper创建后添加以下代码
// 移除所有注释
for (int i = 1; i <= reader.getNumberOfPages(); i++) {
    PdfDictionary pageDict = reader.getPageN(i);
    pageDict.remove(PdfName.ANNOTS);
}
// 移除未使用的字体
reader.removeUnusedFonts();

二、PDFBox 优化方案

如果你切换使用PDFBox,也可以通过以下方式压缩:

1. 基础文档压缩

启用内容流和交叉引用流的压缩:

try (PDDocument document = PDDocument.load(new File(inputPdfPath))) {
    SaveOptions saveOptions = new SaveOptions();
    // 压缩内容流
    saveOptions.setCompressContentStreams(true);
    // 压缩交叉引用流
    saveOptions.setCompressXrefStreams(true);
    // 设置最高压缩级别
    saveOptions.setCompressionLevel(9);
    
    document.save(outputPdfPath, saveOptions);
} catch (IOException e) {
    e.printStackTrace();
}

2. 图片压缩与优化

遍历并压缩文档中的图片:

try (PDDocument document = PDDocument.load(new File(inputPdfPath))) {
    for (PDPage page : document.getPages()) {
        PDResources resources = page.getResources();
        for (COSName name : resources.getXObjectNames()) {
            PDXObject xobject = resources.getXObject(name);
            if (xobject instanceof PDImageXObject) {
                PDImageXObject image = (PDImageXObject) xobject;
                
                // 降低分辨率至150DPI(适合打印/屏幕显示)
                image.setDPI(150, 150);
                
                // 将PNG转为JPEG压缩(质量70%)
                if ("png".equalsIgnoreCase(image.getSuffix())) {
                    BufferedImage bufferedImage = image.getImage();
                    ByteArrayOutputStream baos = new ByteArrayOutputStream();
                    ImageWriteParam param = ImageIO.getImageWritersByFormatName("JPEG").next().getDefaultWriteParam();
                    param.setCompressionMode(ImageWriteParam.MODE_EXPLICIT);
                    param.setCompressionQuality(0.7f);
                    
                    ImageIO.write(bufferedImage, "JPEG", new javax.imageio.stream.MemoryCacheImageOutputStream(baos));
                    PDImageXObject compressedImage = PDImageXObject.createFromByteArray(document, baos.toByteArray(), "JPEG");
                    resources.putXObject(name, compressedImage);
                }
            }
        }
    }
    
    // 清理未使用的资源
    document.cleanUp();
    
    SaveOptions saveOptions = new SaveOptions();
    saveOptions.setCompressContentStreams(true);
    saveOptions.setCompressionLevel(9);
    document.save(outputPdfPath, saveOptions);
} catch (IOException e) {
    e.printStackTrace();
}

3. 移除冗余元素

try (PDDocument document = PDDocument.load(new File(inputPdfPath))) {
    // 移除表单域
    document.getDocumentCatalog().setAcroForm(null);
    // 移除所有页面注释
    for (PDPage page : document.getPages()) {
        page.setAnnotations(new ArrayList<>());
    }
    // 清除元数据(如作者、标题)
    PDDocumentInformation info = document.getDocumentInformation();
    info.setAuthor(null);
    info.setTitle(null);
    info.setSubject(null);
    
    document.cleanUp();
    document.save(outputPdfPath);
} catch (IOException e) {
    e.printStackTrace();
}

内容的提问来源于stack exchange,提问作者Kishor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 22:18:14