You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PDFsharp转换XML的CDATA节点为PDF时文档损坏求助

问题:用PDFsharp将XML中CDATA节点内容转PDF后文件损坏

我尝试通过PDFsharp读取XML文档里的CDATA节点内容,将其转换为PDF文档,但生成的PDF文件始终损坏无法正常打开。

测试代码

Console.WriteLine("Read input file.");

String xmlFile = File.ReadAllText("input.xml");
XDocument xmlDoc = XDocument.Parse(xmlFile);
XElement node1Element = xmlDoc.Descendants("tekening").FirstOrDefault();
string cdataValue = node1Element.Element("tekening_data").Value;

Console.Write(cdataValue);
Console.WriteLine();

byte[] bytes = Encoding.UTF8.GetBytes(cdataValue);

Console.WriteLine("Create PDF from input file.");
using (PdfDocument InputDocument = PdfReader.Open(new MemoryStream(bytes), PdfDocumentOpenMode.Import))
{
    Console.WriteLine("Save as PDF file.");
    InputDocument.Save("output.pdf");
}

Console.WriteLine("Ready!");
Console.ReadLine();

XML输入

<tekening>
    <tekening_data><![CDATA[%%Title: Platten-Skizze von AB: 20133/061123 Pos.: <0048|00> WB65
%%From: CNCBD
%%Creator: GAV Vers 5.2.21552
%%CreationDate: 2023-04-20 11:14:56
%Driver: PcsPrint 5.0.0
% *** Begin embedded object, PcsPrint 5.0.0 ***
/SaveMark_Embed save def
% save status
%%BeginProlog
%%
%% Standard prolog
%% (c) Copyright PCS Strakeljahn GmbH & Co. KG 2011
%%
/BIND {bind def} bind def /_SNAP {transform round exch round exch itransform} BIND /_DSNAP {dtransform round exch round exch
idtransform} BIND /LINE {_SNAP lineto} BIND /MOVE {_SNAP moveto} BIND /RLINE {_DSNAP rlineto} BIND 

.... snip ...

08WL6) CENTER GR
% <------------- END APL-Core
/DS {1.0000 mul} BIND GR
% Ende Platten-Skizze von AB: 20133/061123 Pos.: <0048|00> WB65
SaveMark_Embed restore
% *** End embedded object ***]]></tekening_data>
</tekening>

已尝试的操作

我已经试过以下两种操作,但问题依旧:

  • 修剪CDATA内容中的换行符
  • 移除CDATA节点里附带的头部内容(如下所示):
650 800 translate 180 rotate
<< /Orientation 2 >> setpagedevice
50 10 translate .95 .95 scale

解决方案

核心问题是:CDATA里的内容不是完整的PDF文件,而是PDF内部的PostScript片段(嵌入式打印作业内容),PDFsharp的PdfReader.Open只能解析完整的PDF文件结构,自然无法识别这种片段,导致生成损坏的文件。

正确的处理方式是借助Ghostscript将PostScript片段转换为完整PDF,再用PDFsharp处理,修改后的代码示例:

Console.WriteLine("Read input file.");

String xmlFile = File.ReadAllText("input.xml");
XDocument xmlDoc = XDocument.Parse(xmlFile);
XElement node1Element = xmlDoc.Descendants("tekening").FirstOrDefault();
string cdataValue = node1Element.Element("tekening_data").Value;

// 提取PostScript核心内容,去除前后标记注释
string psContent = cdataValue
    .Replace("% *** Begin embedded object, PcsPrint 5.0.0 ***", "")
    .Replace("% *** End embedded object ***", "")
    .Trim();

Console.WriteLine("Create new PDF document.");
using (PdfDocument document = new PdfDocument())
{
    // 调用Ghostscript将PostScript转换为临时PDF
    var psi = new ProcessStartInfo
    {
        FileName = @"C:\Program Files\gs\gs10.02.0\bin\gswin64c.exe", // 替换为你的Ghostscript路径
        Arguments = "-sDEVICE=pdfwrite -o temp.pdf -f -",
        RedirectStandardInput = true,
        UseShellExecute = false,
        CreateNoWindow = true
    };
    
    using (var process = Process.Start(psi))
    {
        process.StandardInput.Write(psContent);
        process.StandardInput.Close();
        process.WaitForExit();
    }
    
    // 导入临时PDF到新文档
    using (var tempPdf = PdfReader.Open("temp.pdf", PdfDocumentOpenMode.Import))
    {
        foreach (PdfPage tempPage in tempPdf.Pages)
        {
            document.AddPage(tempPage);
        }
    }
    
    // 删除临时文件
    File.Delete("temp.pdf");
    
    Console.WriteLine("Save as PDF file.");
    document.Save("output.pdf");
}

Console.WriteLine("Ready!");
Console.ReadLine();

补充说明:

  • PDFsharp本身不支持PostScript解析,必须依赖Ghostscript这类工具完成格式转换
  • 确保已安装Ghostscript,并将代码中的执行文件路径替换为你本地的实际路径
  • 若无需额外PDF处理,也可直接用Ghostscript将CDATA内容转成PDF,跳过PDFsharp步骤

内容的提问来源于stack exchange,提问作者user11298127

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 22:44:55