You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DocIO转换HTML到SFDT时因非XHTML格式报错,如何关闭验证?

问题:Primeface RichEditor迁移至Syncfusion WordEditor的HTML格式兼容问题

我正在从旧的Primeface RichEditor切换到Syncfusion WordEditor,使用以下类实现HTML与SFDT的互转:

import java.io.ByteArrayInputStream;
import java.io.InputStream;
import java.nio.charset.StandardCharsets;
import com.syncfusion.docio.FormatType;
import com.syncfusion.docio.WordDocument;
import com.syncfusion.ej2.wordprocessor.WordProcessorHelper;

public class SFDTAdapter {
    
    public static String sfdtToRtf(String sfdt) throws Exception {
        return WordProcessorHelper.save(sfdt, com.syncfusion.ej2.wordprocessor.FormatType.Rtf).toString();
    }
    
    public static String rtfToSfdt(String rtf) throws Exception {
        byte[] bytes = rtf.getBytes(StandardCharsets.UTF_8);
        InputStream stream = new ByteArrayInputStream(bytes);
        WordDocument document = new WordDocument(stream, FormatType.Rtf);
        String sfdt =  WordProcessorHelper.load(document);
        document.close();
        stream.close();
        return sfdt;
    }
    
    
    public static String htmlToSfdt(String html) throws Exception {
        byte[] bytes = html.getBytes(StandardCharsets.UTF_8);
        InputStream stream = new ByteArrayInputStream(bytes);
        WordDocument document = new WordDocument(stream, FormatType.Html);
        String sfdt = WordProcessorHelper.load(document);
        document.close();
        stream.close();
        return sfdt;
    }
    
    public static String sfdtToHtml(String sfdt) throws Exception {
        return WordProcessorHelper.save(sfdt, com.syncfusion.ej2.wordprocessor.FormatType.Html).toString();
    }
}

处理旧数据时触发错误:The element type 'br' must be terminated by the matching end-tag。已知DocIO要求内容符合XHTML 1格式,请问能否让DocIO忽略错误或跳过格式验证?


解决方案

Syncfusion DocIO本身没有提供直接关闭HTML格式验证的选项,因为它依赖严格的XHTML解析规则,但可以通过以下两种方式解决该问题:

1. 预处理旧HTML,转换为合规XHTML

旧数据中像<br>这类未闭合标签是报错的核心原因,你可以用HTML解析库自动修正格式,推荐使用Jsoup(轻量级且易用):

步骤1:添加预处理方法

import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;

private static String convertToValidXhtml(String rawHtml) {
    // Jsoup自动解析松散HTML并输出标准XHTML
    Document doc = Jsoup.parse(rawHtml);
    doc.outputSettings().syntax(Document.OutputSettings.Syntax.xml);
    return doc.html();
}

步骤2:修改htmlToSfdt方法,先处理再导入

public static String htmlToSfdt(String html) throws Exception {
    // 先将旧HTML转换为合规XHTML
    String validXhtml = convertToValidXhtml(html);
    byte[] bytes = validXhtml.getBytes(StandardCharsets.UTF_8);
    
    // 使用try-with-resources自动关闭资源,避免手动遗漏
    try (InputStream stream = new ByteArrayInputStream(bytes);
         WordDocument document = new WordDocument(stream, FormatType.Html)) {
        return WordProcessorHelper.load(document);
    }
}

2. 配合HtmlImportOptions忽略非致命解析错误

虽然无法完全关闭验证,但可以开启IgnoreParseErrors选项来忽略非致命的解析问题,配合预处理使用能进一步提升兼容性:

import com.syncfusion.docio.HtmlImportOptions;

public static String htmlToSfdt(String html) throws Exception {
    String validXhtml = convertToValidXhtml(html);
    byte[] bytes = validXhtml.getBytes(StandardCharsets.UTF_8);
    
    try (InputStream stream = new ByteArrayInputStream(bytes)) {
        HtmlImportOptions importOpts = new HtmlImportOptions();
        importOpts.setIgnoreParseErrors(true); // 忽略非致命解析错误
        
        try (WordDocument document = new WordDocument(stream, FormatType.Html, importOpts)) {
            return WordProcessorHelper.load(document);
        }
    }
}

内容的提问来源于stack exchange,提问作者Eduardo Roque

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 10:18:11