DocIO转换HTML到SFDT时因非XHTML格式报错,如何关闭验证?
问题:Primeface RichEditor迁移至Syncfusion WordEditor的HTML格式兼容问题
我正在从旧的Primeface RichEditor切换到Syncfusion WordEditor,使用以下类实现HTML与SFDT的互转:
import java.io.ByteArrayInputStream; import java.io.InputStream; import java.nio.charset.StandardCharsets; import com.syncfusion.docio.FormatType; import com.syncfusion.docio.WordDocument; import com.syncfusion.ej2.wordprocessor.WordProcessorHelper; public class SFDTAdapter { public static String sfdtToRtf(String sfdt) throws Exception { return WordProcessorHelper.save(sfdt, com.syncfusion.ej2.wordprocessor.FormatType.Rtf).toString(); } public static String rtfToSfdt(String rtf) throws Exception { byte[] bytes = rtf.getBytes(StandardCharsets.UTF_8); InputStream stream = new ByteArrayInputStream(bytes); WordDocument document = new WordDocument(stream, FormatType.Rtf); String sfdt = WordProcessorHelper.load(document); document.close(); stream.close(); return sfdt; } public static String htmlToSfdt(String html) throws Exception { byte[] bytes = html.getBytes(StandardCharsets.UTF_8); InputStream stream = new ByteArrayInputStream(bytes); WordDocument document = new WordDocument(stream, FormatType.Html); String sfdt = WordProcessorHelper.load(document); document.close(); stream.close(); return sfdt; } public static String sfdtToHtml(String sfdt) throws Exception { return WordProcessorHelper.save(sfdt, com.syncfusion.ej2.wordprocessor.FormatType.Html).toString(); } }
处理旧数据时触发错误:The element type 'br' must be terminated by the matching end-tag。已知DocIO要求内容符合XHTML 1格式,请问能否让DocIO忽略错误或跳过格式验证?
解决方案
Syncfusion DocIO本身没有提供直接关闭HTML格式验证的选项,因为它依赖严格的XHTML解析规则,但可以通过以下两种方式解决该问题:
1. 预处理旧HTML,转换为合规XHTML
旧数据中像<br>这类未闭合标签是报错的核心原因,你可以用HTML解析库自动修正格式,推荐使用Jsoup(轻量级且易用):
步骤1:添加预处理方法
import org.jsoup.Jsoup; import org.jsoup.nodes.Document; private static String convertToValidXhtml(String rawHtml) { // Jsoup自动解析松散HTML并输出标准XHTML Document doc = Jsoup.parse(rawHtml); doc.outputSettings().syntax(Document.OutputSettings.Syntax.xml); return doc.html(); }
步骤2:修改htmlToSfdt方法,先处理再导入
public static String htmlToSfdt(String html) throws Exception { // 先将旧HTML转换为合规XHTML String validXhtml = convertToValidXhtml(html); byte[] bytes = validXhtml.getBytes(StandardCharsets.UTF_8); // 使用try-with-resources自动关闭资源,避免手动遗漏 try (InputStream stream = new ByteArrayInputStream(bytes); WordDocument document = new WordDocument(stream, FormatType.Html)) { return WordProcessorHelper.load(document); } }
2. 配合HtmlImportOptions忽略非致命解析错误
虽然无法完全关闭验证,但可以开启IgnoreParseErrors选项来忽略非致命的解析问题,配合预处理使用能进一步提升兼容性:
import com.syncfusion.docio.HtmlImportOptions; public static String htmlToSfdt(String html) throws Exception { String validXhtml = convertToValidXhtml(html); byte[] bytes = validXhtml.getBytes(StandardCharsets.UTF_8); try (InputStream stream = new ByteArrayInputStream(bytes)) { HtmlImportOptions importOpts = new HtmlImportOptions(); importOpts.setIgnoreParseErrors(true); // 忽略非致命解析错误 try (WordDocument document = new WordDocument(stream, FormatType.Html, importOpts)) { return WordProcessorHelper.load(document); } } }
内容的提问来源于stack exchange,提问作者Eduardo Roque
相关产品推荐
相关产品推荐

