You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用iTextSharp 5将HTML元素转换为PDF/A格式的实现方法

解决iTextSharp生成PDF/A时HTML标签无法渲染格式的问题

Absolutely, you can use XMLWorkerHelper to parse HTML tags (like <b>, <i>) into properly formatted content in your PDF/A document. Your current code treats escaped HTML entities as plain text, which is why the formatting isn't applied. Let's walk through how to adjust your implementation:

Key Changes Needed

  • Stop escaping HTML tags: Use raw tags like <b>bold text</b> instead of &lt;b&gt;bold text&lt;/b&gt;
  • Use XMLWorkerHelper.ParseXHtml() to convert HTML content directly into iText elements that respect formatting
  • Ensure PDF/A requirements (font embedding, output intent) are still maintained while using XMLWorker

Modified Full Code

using System;
using System.IO;
using System.Text;
using iTextSharp.text;
using iTextSharp.text.pdf;
using iTextSharp.tool.xml;
using iTextSharp.tool.xml.pipeline.html;

namespace itext_html {
    class Program {
        static void Main(string[] args) {
            // Initialize font (embedded, required for PDF/A)
            var arialPath = Path.Combine(Environment.GetEnvironmentVariable("SystemRoot"), "fonts", "arial.ttf");
            var baseFontArial = BaseFont.CreateFont(arialPath, BaseFont.IDENTITY_H, BaseFont.EMBEDDED);
            var arial7 = new Font(baseFontArial, 7);

            // Set up PDF/A document and writer
            var outputPath = @"c:\temp\itext\itext_pdfa_html.pdf";
            using (var stream = new FileStream(outputPath, FileMode.Create))
            using (var doc = new Document(PageSize.A4))
            {
                var pdfAWriter = PdfAWriter.GetInstance(doc, stream, PdfAConformanceLevel.PDF_A_1B);
                doc.Open();

                // Set PDF/A output intent (required for color profile compliance)
                var icc = ICC_Profile.GetInstance(@"C:\Temp\itext\srgb.profile");
                pdfAWriter.SetOutputIntents("Custom", "", "http://www.color.org", "sRGB IEC61966-2.1", icc);

                // --- 1. Add formatted HTML content to a PdfPTable ---
                var table = new PdfPTable(2);
                table.SetWidths(new int[] { 65, 35 });

                // HTML content for first cell
                var htmlText1 = "Hello World <b>bold text</b>";
                var cell1 = new PdfPCell();
                // Use XMLWorker to parse HTML into the cell
                using (var stringReader = new StringReader(htmlText1))
                {
                    // Configure font provider to use our embedded Arial font
                    var fontProvider = new XMLWorkerFontProvider();
                    fontProvider.Register(arialPath);
                    var htmlPipelineContext = new HtmlPipelineContext(new CssAppliersImpl(fontProvider));
                    htmlPipelineContext.SetTagFactory(Tags.GetHtmlTagProcessorFactory());

                    var pdfWriterPipeline = new PdfWriterPipeline(doc, pdfAWriter);
                    var htmlPipeline = new HtmlPipeline(htmlPipelineContext, pdfWriterPipeline);
                    var cssResolver = XMLWorkerHelper.GetInstance().GetDefaultCssResolver(true);
                    var pipeline = new CssResolverPipeline(cssResolver, htmlPipeline);

                    var worker = new XMLWorker(pipeline, true);
                    var parser = new XMLParser(worker);
                    parser.Parse(stringReader);

                    // Add parsed elements to the cell
                    foreach (var element in doc.GetElements())
                    {
                        cell1.AddElement(element);
                    }
                    // Clear the document's elements to avoid duplication
                    doc.Clear();
                }
                table.AddCell(cell1);

                // HTML content for second cell
                var htmlText2 = "Another text <i>italic text</i>";
                var cell2 = new PdfPCell();
                using (var stringReader = new StringReader(htmlText2))
                {
                    var fontProvider = new XMLWorkerFontProvider();
                    fontProvider.Register(arialPath);
                    var htmlPipelineContext = new HtmlPipelineContext(new CssAppliersImpl(fontProvider));
                    htmlPipelineContext.SetTagFactory(Tags.GetHtmlTagProcessorFactory());

                    var pdfWriterPipeline = new PdfWriterPipeline(doc, pdfAWriter);
                    var htmlPipeline = new HtmlPipeline(htmlPipelineContext, pdfWriterPipeline);
                    var cssResolver = XMLWorkerHelper.GetInstance().GetDefaultCssResolver(true);
                    var pipeline = new CssResolverPipeline(cssResolver, htmlPipeline);

                    var worker = new XMLWorker(pipeline, true);
                    var parser = new XMLParser(worker);
                    parser.Parse(stringReader);

                    foreach (var element in doc.GetElements())
                    {
                        cell2.AddElement(element);
                    }
                    doc.Clear();
                }
                table.AddCell(cell2);

                doc.Add(table);

                // --- 2. Add standalone formatted HTML text ---
                var standaloneHtml = "Standalone text: <b>Bold</b> and <i>Italic</i> example";
                using (var stringReader = new StringReader(standaloneHtml))
                {
                    var fontProvider = new XMLWorkerFontProvider();
                    fontProvider.Register(arialPath);
                    XMLWorkerHelper.GetInstance().ParseXHtml(pdfAWriter, doc, stringReader, null, Encoding.UTF8, fontProvider);
                }

                // Finalize PDF/A document
                pdfAWriter.CreateXmpMetadata();
                doc.Close();
            }
        }
    }
}

Important Notes

  • Font Embedding: We register the Arial font with XMLWorkerFontProvider to ensure it's embedded in the PDF/A document (a mandatory requirement for PDF/A compliance).
  • PDF/A Compliance: We retain the output intent setup (ICC profile) to meet PDF/A color standards.
  • Cell Content Handling: When adding HTML to a PdfPCell, we parse the HTML first, collect the generated elements, add them to the cell, then clear the document's element list to prevent duplicate content.
  • Standalone HTML: For direct document content, XMLWorkerHelper.ParseXHtml() simplifies the process by handling the pipeline setup for you.

内容的提问来源于stack exchange,提问作者Batar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:10:14