使用PDFsharp转换XML的CDATA节点为PDF时文档损坏求助
问题:用PDFsharp将XML中CDATA节点内容转PDF后文件损坏
我尝试通过PDFsharp读取XML文档里的CDATA节点内容,将其转换为PDF文档,但生成的PDF文件始终损坏无法正常打开。
测试代码
Console.WriteLine("Read input file."); String xmlFile = File.ReadAllText("input.xml"); XDocument xmlDoc = XDocument.Parse(xmlFile); XElement node1Element = xmlDoc.Descendants("tekening").FirstOrDefault(); string cdataValue = node1Element.Element("tekening_data").Value; Console.Write(cdataValue); Console.WriteLine(); byte[] bytes = Encoding.UTF8.GetBytes(cdataValue); Console.WriteLine("Create PDF from input file."); using (PdfDocument InputDocument = PdfReader.Open(new MemoryStream(bytes), PdfDocumentOpenMode.Import)) { Console.WriteLine("Save as PDF file."); InputDocument.Save("output.pdf"); } Console.WriteLine("Ready!"); Console.ReadLine();
XML输入
<tekening> <tekening_data><![CDATA[%%Title: Platten-Skizze von AB: 20133/061123 Pos.: <0048|00> WB65 %%From: CNCBD %%Creator: GAV Vers 5.2.21552 %%CreationDate: 2023-04-20 11:14:56 %Driver: PcsPrint 5.0.0 % *** Begin embedded object, PcsPrint 5.0.0 *** /SaveMark_Embed save def % save status %%BeginProlog %% %% Standard prolog %% (c) Copyright PCS Strakeljahn GmbH & Co. KG 2011 %% /BIND {bind def} bind def /_SNAP {transform round exch round exch itransform} BIND /_DSNAP {dtransform round exch round exch idtransform} BIND /LINE {_SNAP lineto} BIND /MOVE {_SNAP moveto} BIND /RLINE {_DSNAP rlineto} BIND .... snip ... 08WL6) CENTER GR % <------------- END APL-Core /DS {1.0000 mul} BIND GR % Ende Platten-Skizze von AB: 20133/061123 Pos.: <0048|00> WB65 SaveMark_Embed restore % *** End embedded object ***]]></tekening_data> </tekening>
已尝试的操作
我已经试过以下两种操作,但问题依旧:
- 修剪CDATA内容中的换行符
- 移除CDATA节点里附带的头部内容(如下所示):
650 800 translate 180 rotate << /Orientation 2 >> setpagedevice 50 10 translate .95 .95 scale
解决方案
核心问题是:CDATA里的内容不是完整的PDF文件,而是PDF内部的PostScript片段(嵌入式打印作业内容),PDFsharp的PdfReader.Open只能解析完整的PDF文件结构,自然无法识别这种片段,导致生成损坏的文件。
正确的处理方式是借助Ghostscript将PostScript片段转换为完整PDF,再用PDFsharp处理,修改后的代码示例:
Console.WriteLine("Read input file."); String xmlFile = File.ReadAllText("input.xml"); XDocument xmlDoc = XDocument.Parse(xmlFile); XElement node1Element = xmlDoc.Descendants("tekening").FirstOrDefault(); string cdataValue = node1Element.Element("tekening_data").Value; // 提取PostScript核心内容,去除前后标记注释 string psContent = cdataValue .Replace("% *** Begin embedded object, PcsPrint 5.0.0 ***", "") .Replace("% *** End embedded object ***", "") .Trim(); Console.WriteLine("Create new PDF document."); using (PdfDocument document = new PdfDocument()) { // 调用Ghostscript将PostScript转换为临时PDF var psi = new ProcessStartInfo { FileName = @"C:\Program Files\gs\gs10.02.0\bin\gswin64c.exe", // 替换为你的Ghostscript路径 Arguments = "-sDEVICE=pdfwrite -o temp.pdf -f -", RedirectStandardInput = true, UseShellExecute = false, CreateNoWindow = true }; using (var process = Process.Start(psi)) { process.StandardInput.Write(psContent); process.StandardInput.Close(); process.WaitForExit(); } // 导入临时PDF到新文档 using (var tempPdf = PdfReader.Open("temp.pdf", PdfDocumentOpenMode.Import)) { foreach (PdfPage tempPage in tempPdf.Pages) { document.AddPage(tempPage); } } // 删除临时文件 File.Delete("temp.pdf"); Console.WriteLine("Save as PDF file."); document.Save("output.pdf"); } Console.WriteLine("Ready!"); Console.ReadLine();
补充说明:
- PDFsharp本身不支持PostScript解析,必须依赖Ghostscript这类工具完成格式转换
- 确保已安装Ghostscript,并将代码中的执行文件路径替换为你本地的实际路径
- 若无需额外PDF处理,也可直接用Ghostscript将CDATA内容转成PDF,跳过PDFsharp步骤
内容的提问来源于stack exchange,提问作者user11298127
相关产品推荐
相关产品推荐

