使用iTextSharp 5将HTML元素转换为PDF/A格式的实现方法
解决iTextSharp生成PDF/A时HTML标签无法渲染格式的问题
Absolutely, you can use XMLWorkerHelper to parse HTML tags (like <b>, <i>) into properly formatted content in your PDF/A document. Your current code treats escaped HTML entities as plain text, which is why the formatting isn't applied. Let's walk through how to adjust your implementation:
Key Changes Needed
- Stop escaping HTML tags: Use raw tags like
<b>bold text</b>instead of<b>bold text</b> - Use
XMLWorkerHelper.ParseXHtml()to convert HTML content directly into iText elements that respect formatting - Ensure PDF/A requirements (font embedding, output intent) are still maintained while using XMLWorker
Modified Full Code
using System; using System.IO; using System.Text; using iTextSharp.text; using iTextSharp.text.pdf; using iTextSharp.tool.xml; using iTextSharp.tool.xml.pipeline.html; namespace itext_html { class Program { static void Main(string[] args) { // Initialize font (embedded, required for PDF/A) var arialPath = Path.Combine(Environment.GetEnvironmentVariable("SystemRoot"), "fonts", "arial.ttf"); var baseFontArial = BaseFont.CreateFont(arialPath, BaseFont.IDENTITY_H, BaseFont.EMBEDDED); var arial7 = new Font(baseFontArial, 7); // Set up PDF/A document and writer var outputPath = @"c:\temp\itext\itext_pdfa_html.pdf"; using (var stream = new FileStream(outputPath, FileMode.Create)) using (var doc = new Document(PageSize.A4)) { var pdfAWriter = PdfAWriter.GetInstance(doc, stream, PdfAConformanceLevel.PDF_A_1B); doc.Open(); // Set PDF/A output intent (required for color profile compliance) var icc = ICC_Profile.GetInstance(@"C:\Temp\itext\srgb.profile"); pdfAWriter.SetOutputIntents("Custom", "", "http://www.color.org", "sRGB IEC61966-2.1", icc); // --- 1. Add formatted HTML content to a PdfPTable --- var table = new PdfPTable(2); table.SetWidths(new int[] { 65, 35 }); // HTML content for first cell var htmlText1 = "Hello World <b>bold text</b>"; var cell1 = new PdfPCell(); // Use XMLWorker to parse HTML into the cell using (var stringReader = new StringReader(htmlText1)) { // Configure font provider to use our embedded Arial font var fontProvider = new XMLWorkerFontProvider(); fontProvider.Register(arialPath); var htmlPipelineContext = new HtmlPipelineContext(new CssAppliersImpl(fontProvider)); htmlPipelineContext.SetTagFactory(Tags.GetHtmlTagProcessorFactory()); var pdfWriterPipeline = new PdfWriterPipeline(doc, pdfAWriter); var htmlPipeline = new HtmlPipeline(htmlPipelineContext, pdfWriterPipeline); var cssResolver = XMLWorkerHelper.GetInstance().GetDefaultCssResolver(true); var pipeline = new CssResolverPipeline(cssResolver, htmlPipeline); var worker = new XMLWorker(pipeline, true); var parser = new XMLParser(worker); parser.Parse(stringReader); // Add parsed elements to the cell foreach (var element in doc.GetElements()) { cell1.AddElement(element); } // Clear the document's elements to avoid duplication doc.Clear(); } table.AddCell(cell1); // HTML content for second cell var htmlText2 = "Another text <i>italic text</i>"; var cell2 = new PdfPCell(); using (var stringReader = new StringReader(htmlText2)) { var fontProvider = new XMLWorkerFontProvider(); fontProvider.Register(arialPath); var htmlPipelineContext = new HtmlPipelineContext(new CssAppliersImpl(fontProvider)); htmlPipelineContext.SetTagFactory(Tags.GetHtmlTagProcessorFactory()); var pdfWriterPipeline = new PdfWriterPipeline(doc, pdfAWriter); var htmlPipeline = new HtmlPipeline(htmlPipelineContext, pdfWriterPipeline); var cssResolver = XMLWorkerHelper.GetInstance().GetDefaultCssResolver(true); var pipeline = new CssResolverPipeline(cssResolver, htmlPipeline); var worker = new XMLWorker(pipeline, true); var parser = new XMLParser(worker); parser.Parse(stringReader); foreach (var element in doc.GetElements()) { cell2.AddElement(element); } doc.Clear(); } table.AddCell(cell2); doc.Add(table); // --- 2. Add standalone formatted HTML text --- var standaloneHtml = "Standalone text: <b>Bold</b> and <i>Italic</i> example"; using (var stringReader = new StringReader(standaloneHtml)) { var fontProvider = new XMLWorkerFontProvider(); fontProvider.Register(arialPath); XMLWorkerHelper.GetInstance().ParseXHtml(pdfAWriter, doc, stringReader, null, Encoding.UTF8, fontProvider); } // Finalize PDF/A document pdfAWriter.CreateXmpMetadata(); doc.Close(); } } } }
Important Notes
- Font Embedding: We register the Arial font with
XMLWorkerFontProviderto ensure it's embedded in the PDF/A document (a mandatory requirement for PDF/A compliance). - PDF/A Compliance: We retain the output intent setup (ICC profile) to meet PDF/A color standards.
- Cell Content Handling: When adding HTML to a
PdfPCell, we parse the HTML first, collect the generated elements, add them to the cell, then clear the document's element list to prevent duplicate content. - Standalone HTML: For direct document content,
XMLWorkerHelper.ParseXHtml()simplifies the process by handling the pipeline setup for you.
内容的提问来源于stack exchange,提问作者Batar
相关产品推荐
相关产品推荐

