You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Java压缩PDF文件大小:将15MB+PDF压缩至最大10MB

解决PDF压缩至10MB以内的方案

一、定位体积超标的核心原因

你的PDF体积超标大概率是图片资源未优化导致的——从代码看你添加了signImg图片,模板PDF本身也可能包含高分辨率图片,这些是体积的主要贡献者。你之前用iText的setFullCompression仅处理文档结构压缩,对图片这类大资源完全无效。

二、基于Apache PDFBox的适配优化(和现有代码兼容)

因为你用PDFBox创建PDF,优先用PDFBox原生工具优化,适配性更好:

1. 提前压缩签名图片

在添加signImg前,先降低图片质量或分辨率,减少体积:

// 对签名图片进行JPEG压缩,调整质量参数平衡体积与清晰度
BufferedImage originalImage = ImageIO.read(new ByteArrayInputStream(signImg));
ByteArrayOutputStream compressedImgStream = new ByteArrayOutputStream();

ImageWriter writer = ImageIO.getImageWritersByFormatName("JPEG").next();
ImageWriteParam param = writer.getDefaultWriteParam();
param.setCompressionMode(ImageWriteParam.MODE_EXPLICIT);
param.setCompressionQuality(0.6f); // 0.1-1.0,值越小体积越小

writer.setOutput(new MemoryCacheImageOutputStream(compressedImgStream));
writer.write(null, new IIOImage(originalImage, null, null), param);
writer.dispose();

byte[] compressedSignImg = compressedImgStream.toByteArray();

// 使用压缩后的图片创建PDImageXObject
PDImageXObject pdImage = PDImageXObject.createFromByteArray(template, compressedSignImg, "sign.jpg");

2. 用PDFBox内置工具整体优化PDF

生成pdfTemplate后,使用PDFOptimizer自动处理图片压缩、冗余资源清理:

ByteArrayOutputStream optimizedStream = new ByteArrayOutputStream();
PDFOptimizer optimizer = new PDFOptimizer();

// 设置优化参数
OptimizationOptions options = new OptimizationOptions();
options.setImageQuality(0.6f); // 全局图片压缩质量
options.setMaxImageSize(72); // 将图片分辨率限制为72dpi(适合屏幕显示)

optimizer.optimize(pdfTemplate, optimizedStream);
pdfTemplate.close();

// 后续可将optimizedStream写入文件或返回

3. 合并模板时共享资源避免冗余

合并多个PDDocument时,重复的字体、图片资源会增加体积,可统一共享资源字典:

PDDocument pdfTemplate = new PDDocument();
PDResources sharedResources = new PDResources();
pdfTemplate.getDocumentCatalog().setResources(sharedResources);

for (PDDocument template : templateList) {
    for (PDPage page : template.getPages()) {
        // 处理签名图片逻辑...
        
        // 合并页面资源到共享字典,避免重复存储
        PDResources pageResources = page.getResources();
        for (COSName name : pageResources.getXObjectNames()) {
            if (!sharedResources.containsKey(name)) {
                sharedResources.put(name, pageResources.getXObject(name));
            }
        }
        pdfTemplate.addPage(page);
    }
    template.close(); // 及时关闭模板文档释放资源
}

三、补充优化策略(若上述方法仍不达标)

  • 清理冗余元数据:删除PDF中不必要的作者、创建日期等元数据:
    PDDocumentInformation info = pdfTemplate.getDocumentInformation();
    info.setAuthor(null);
    info.setCreator(null);
    info.setKeywords(null);
    info.setTitle(null);
    
  • 降低PDF版本:将PDF版本从1.7降至1.5,部分压缩算法更高效:
    pdfTemplate.setVersion(1.5f);
    
  • 调用Ghostscript命令行优化:如果代码层面优化后仍超标,可通过Java执行Ghostscript命令(需提前安装Ghostscript):
    String[] cmd = {
        "gs",
        "-sDEVICE=pdfwrite",
        "-dCompatibilityLevel=1.5",
        "-dPDFSETTINGS=/screen", // /screen为屏幕优化模式,体积最小;/ebook为平衡模式
        "-dNOPAUSE",
        "-dBATCH",
        "-sOutputFile=output.pdf",
        "input.pdf"
    };
    Runtime.getRuntime().exec(cmd);
    

四、为什么你之前的iText方法无效?

你用的reader.setPageContent(i + 1, reader.getPageContent(i + 1))只是重复设置页面内容,没有实际压缩逻辑;setFullCompression仅对未压缩的文档结构有效,若原PDF已压缩则无作用。另外iText与PDFBox的PDF处理逻辑不兼容,跨库处理容易出现无效情况,建议统一用PDFBox完成创建与压缩流程。

内容的提问来源于stack exchange,提问作者Preeti Joshi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 02:38:10