Docx转HTML无法识别标题与列表?跨库兼容问题排查
问题描述
我用com.spire.doc生成Docx文件,现在要基于org.apache.poi和org.jsoup.nodes实现Docx转HTML功能。转换器能正常运行,但无法识别文档中的标题和列表,添加标题判断条件后也不生效。不确定是跨库兼容问题,还是缺少正确的标题、列表检测方法,以下是相关内容,寻求解决方案:
输入Docx内容
C++ <-- header C++ is a general purpose, high-level programming language developed by Sun Microsystems. The C++ programming language was developed by a small team of engineers, known as the Green Team, who initiated the language in 1991. - Chapter 1 - Chapter 2 - Chapter 3
转换器代码
package lab2.converter; import java.io.*; import java.util.*; import org.apache.poi.xwpf.usermodel.*; import org.jsoup.nodes.Document; import org.jsoup.nodes.Element; public class HtmlConverter implements Converter{ @Override public void convert() throws IOException { try { // 1. 用Apache POI读取Word文档 FileInputStream fis = new FileInputStream("output/Find&Replace.docx"); XWPFDocument document = new XWPFDocument(fis); // 2. 用jsoup转成HTML Document htmlDoc = new Document(""); Element html = htmlDoc.appendElement("html"); Element body = html.appendElement("body"); List<XWPFParagraph> paragraphs = document.getParagraphs(); for (XWPFParagraph paragraph : paragraphs) { Element p = body.appendElement("p"); String text = paragraph.getText(); // 检测标题和列表的判断条件 // 未实现 // 把文本添加到段落 p.appendText(text); } // 3. 把HTML写入文件 try (PrintWriter out = new PrintWriter("output/Find&Replace.html")) { out.println(htmlDoc.html()); } System.out.println("转换完成。"); document.close(); } catch (Exception e) { System.err.println("Word转HTML出错:" + e.getMessage()); } } }
转换后的HTML结果
<html> <body> <p>Evaluation Warning: The document was created with Spire.Doc for JAVA.</p> <p>Evaluation Warning: The document was created with Spire.Doc for C++.</p> <p>C++</p> <p>C++ is a general purpose, high-level programming language developed by Sun Microsystems. The C++ programming language was developed by a small team of engineers, known as the Green Team, who initiated the language in 1991.</p> <p>Chapter 1</p> <p>Chapter 2</p> <p>Chapter 3</p> </body> </html>
解决方案
1. 标题识别:通过样式或格式判断
Spire.Doc生成的标题会对应Word内置的标题样式(如Heading 1),可通过XWPFParagraph的样式信息或字体格式识别:
方法一:通过样式ID/名称判断
for (XWPFParagraph paragraph : paragraphs) { String text = paragraph.getText().trim(); // 跳过空段落和Spire评估警告 if (text.isEmpty() || text.startsWith("Evaluation Warning")) { continue; } Element contentElement; // 检查段落样式 String styleId = paragraph.getStyleID(); String styleName = paragraph.getStyle() != null ? paragraph.getStyle().getName() : ""; // 匹配Word内置标题样式 if ((styleId != null && styleId.startsWith("Heading")) || styleName.startsWith("Heading")) { // 根据标题级别生成h1-h6,这里以Heading 1为例对应h1 contentElement = body.appendElement("h1"); } else { contentElement = body.appendElement("p"); } contentElement.appendText(text); }
方法二:通过字体格式辅助判断
如果Spire生成的样式非标准Heading开头,可通过字体大小、加粗属性判断:
if (!paragraph.getRuns().isEmpty()) { XWPFRun firstRun = paragraph.getRuns().get(0); if (firstRun.isBold() && firstRun.getFontSize() >= 16) { contentElement = body.appendElement("h1"); } }
2. 列表识别:检查段落的列表格式
Apache POI可通过段落的getNumPr()属性判断是否为列表项,同时维护当前列表元素避免重复创建:
Element currentUl = null; for (XWPFParagraph paragraph : paragraphs) { String text = paragraph.getText().trim(); if (text.isEmpty() || text.startsWith("Evaluation Warning")) { continue; } // 判断是否为列表项 boolean isListItem = paragraph.getCTP().getPPr() != null && paragraph.getCTP().getPPr().getNumPr() != null; if (isListItem) { if (currentUl == null) { currentUl = body.appendElement("ul"); } // 移除文本开头的"- "标记 currentUl.appendElement("li").appendText(text.replaceFirst("^- ", "")); } else { // 退出列表模式 currentUl = null; // 处理标题或普通段落 Element contentElement; String styleName = paragraph.getStyle() != null ? paragraph.getStyle().getName() : ""; if (styleName.startsWith("Heading")) { contentElement = body.appendElement("h1"); } else { contentElement = body.appendElement("p"); } contentElement.appendText(text); } }
3. 跨库兼容注意事项
Spire.Doc生成的Docx符合OOXML标准,Apache POI可正常解析,但需注意:
- 免费版Spire会插入评估警告段落,需过滤
- 部分自定义样式的ID/名称可能与POI默认识别规则有差异,必要时结合字体、段落格式辅助判断
内容的提问来源于stack exchange,提问作者apo
相关产品推荐
相关产品推荐

