You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用iText 5的XMLWorkerHelper将含HTML标签的字符串写入PDF并保留样式

解决iText 5中XMLWorkerHelper无法完整渲染HTML内容的问题

我完全懂你的困扰——直接用XMLWorkerHelper.getInstance().parseXHtml()确实容易出现只解析标签内内容、忽略其余文本的情况,而HTMLWorker又已经被弃用,总不能抱着过时的API不放对吧?别担心,咱们用iText 5的XMLWorker底层API就能完美解决这个问题,而且完全符合官方推荐的用法。

问题根源

你之前的写法太简化了,parseXHtml的默认实现没有正确初始化HTML处理上下文,导致部分文本节点和标签样式没有被正确解析。咱们需要手动构建XMLWorker的处理管道,确保所有HTML内容都能被正确识别和渲染。

可行实现方案

下面是完整的可运行代码,不仅能正确渲染<sup>上标标签,还能支持更多HTML样式(比如<b>、<i>、<p>等):

import com.itextpdf.text.Document;
import com.itextpdf.text.PageSize;
import com.itextpdf.text.pdf.PdfWriter;
import com.itextpdf.tool.xml.XMLWorkerHelper;
import com.itextpdf.tool.xml.css.StyleAttrCSSResolver;
import com.itextpdf.tool.xml.parser.XMLParser;
import com.itextpdf.tool.xml.pipeline.css.CSSResolver;
import com.itextpdf.tool.xml.pipeline.css.CssResolverPipeline;
import com.itextpdf.tool.xml.pipeline.end.PdfWriterPipeline;
import com.itextpdf.tool.xml.pipeline.html.HtmlPipeline;
import com.itextpdf.tool.xml.pipeline.html.HtmlPipelineContext;

import java.io.FileOutputStream;
import java.io.StringReader;

public class HtmlToPdfRenderer {
    public static void main(String[] args) throws Exception {
        // 1. 初始化PDF文档和Writer
        Document document = new Document(PageSize.A4);
        PdfWriter writer = PdfWriter.getInstance(document, new FileOutputStream("math_question.pdf"));
        document.open();

        // 2. 配置HTML处理上下文,确保标签被正确识别
        HtmlPipelineContext htmlContext = new HtmlPipelineContext(null);
        htmlContext.setTagFactory(com.itextpdf.tool.xml.html.Tags.getHtmlTagProcessorFactory());
        
        // 注册系统字体,避免样式渲染异常(比如上标显示不正常)
        com.itextpdf.text.FontFactory.registerDirectories();

        // 3. 创建CSS解析器,支持自定义样式
        CSSResolver cssResolver = new StyleAttrCSSResolver();
        // 可选:添加自定义CSS来精确控制上标样式
        cssResolver.addCss("sup { font-size: 0.75em; vertical-align: super; margin-left: 0.1em; }", true);

        // 4. 构建XMLWorker处理管道:CSS解析 -> HTML解析 -> PDF输出
        PdfWriterPipeline pdfOutputPipeline = new PdfWriterPipeline(document, writer);
        HtmlPipeline htmlProcessingPipeline = new HtmlPipeline(htmlContext, pdfOutputPipeline);
        CssResolverPipeline cssProcessingPipeline = new CssResolverPipeline(cssResolver, htmlProcessingPipeline);

        // 5. 解析并渲染HTML内容
        String htmlContent = "What is the equation of the line passing through the point (2,-3) and making an angle of -45<sup>2</sup> with the positive X-axis?";
        XMLParser parser = new XMLParser(cssProcessingPipeline);
        parser.parse(new StringReader(htmlContent));

        // 6. 关闭资源
        document.close();
        writer.close();
    }
}

关键说明

  • 手动构建处理管道:通过CssResolverPipeline、HtmlPipeline和PdfWriterPipeline的组合,确保HTML的文本节点和标签都能被完整解析,不会出现内容丢失的情况。
  • 字体注册:调用FontFactory.registerDirectories()注册系统字体,保证上标、特殊字符等样式能正确渲染,避免默认字体缺失导致的异常。
  • 自定义CSS:可以通过cssResolver.addCss()添加自定义样式规则,更精确地控制HTML元素的显示效果,比如调整上标的大小和位置。

这个方案完全基于iText 5的官方推荐API,没有使用已弃用的HTMLWorker,能稳定保留HTML标签对应的样式,适用于各种复杂程度的HTML内容。

内容的提问来源于stack exchange,提问作者Nilendu Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:39:23