You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PDFSharp C#拆分合并PDF后文件体积增大的原因及解决办法

PDF拆分合并后文件体积暴涨的问题及解决方案

问题背景

使用PDFSharp的C#代码将PDF拆分为单页文档,再合并回单PDF后,文件体积从1.96MB增至12.2MB,已确认问题出在拆分合并环节而非中间操作。相关代码如下:

public static List<Stream> SplitPdf(Stream PdfDoc)
{
    System.Text.Encoding.RegisterProvider(System.Text.CodePagesEncodingProvider.Instance);
    List<Stream> outputStreamList = new List<Stream>();
    PdfSharp.Pdf.PdfDocument inputDocument = PdfReader.Open(PdfDoc, PdfDocumentOpenMode.Import);

    for (int idx = 0; idx < inputDocument.PageCount; idx++)
    {
        PdfSharp.Pdf.PdfDocument outputDocument = new PdfSharp.Pdf.PdfDocument();
        outputDocument.Version = inputDocument.Version;
        outputDocument.Info.Title =
          String.Format("Page {0} of {1}", idx + 1, inputDocument.Info.Title);
        outputDocument.Info.Creator = inputDocument.Info.Creator;

        outputDocument.AddPage(inputDocument.Pages[idx]);
        MemoryStream stream = new MemoryStream();
        outputDocument.Save(stream);
        outputStreamList.Add(stream);
    }
    return outputStreamList;
}

public static Stream MergePdfs(List<Stream> PdfFiles)
{
    System.Text.Encoding.RegisterProvider(System.Text.CodePagesEncodingProvider.Instance);
    PdfSharp.Pdf.PdfDocument outputPDFDocument = new PdfSharp.Pdf.PdfDocument();
    foreach (Stream pdfFile in PdfFiles)
    {
        PdfSharp.Pdf.PdfDocument inputPDFDocument = PdfReader.Open(pdfFile, PdfDocumentOpenMode.Import);
        outputPDFDocument.Version = inputPDFDocument.Version;
        foreach (PdfSharp.Pdf.PdfPage page in inputPDFDocument.Pages)
        {
            outputPDFDocument.AddPage(page);
        }
    }
    Stream compiledPdfStream = new MemoryStream();
    outputPDFDocument.Save(compiledPdfStream);
    return compiledPdfStream;
}

问题

  1. 为何会出现体积暴涨的情况?
  2. 是否有开源C#库能实现拆分合并后保持原文件大小?

解答

1. 体积暴涨的原因

  • 共享资源重复嵌入:原PDF中的共享资源(如字体、图片、表单模板等)会被每个拆分出的单页PDF单独复制一份,合并时这些重复资源全部保留,导致总体积剧增。
  • PDFSharp的资源处理逻辑:使用PdfDocumentOpenMode.Import导入页面时,PDFSharp会将页面依赖的所有资源完整复制到新文档,而非引用原文档的共享资源池,拆分后的单页文档都带有独立资源副本,合并后无法自动去重。
  • 压缩策略不一致:原PDF可能采用了高效压缩算法(如JPEG2000、优化后的Flate压缩),而PDFSharp保存文档时默认压缩设置较宽松,未匹配原文档的压缩参数。

2. 保持原文件大小的开源C#库解决方案

  • iText 7 .NET:支持PDF资源的共享引用,拆分时可保留原文档的资源池,合并时不会重复嵌入资源,同时可配置与原文档一致的压缩参数,确保体积接近原文件。
  • PdfPig:基于Apache PDFBox的.NET实现,高效处理PDF拆分与合并,默认逻辑会尽量保留原文档的结构和压缩特性,避免资源重复。
  • QuestPDF:虽主打PDF生成,但支持PDF合并操作,底层优化了资源处理逻辑,能避免不必要的资源重复,维持文件体积稳定。

使用这些库时,需采用“导入页面而非强制复制资源”的模式,并匹配原文档的压缩设置,才能最大程度保持原文件体积。


内容的提问来源于stack exchange,提问作者Abdullah Ahmed khan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 04:35:36