You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Docx转HTML无法识别标题与列表?跨库兼容问题排查

问题描述

我用com.spire.doc生成Docx文件,现在要基于org.apache.poi和org.jsoup.nodes实现Docx转HTML功能。转换器能正常运行,但无法识别文档中的标题和列表,添加标题判断条件后也不生效。不确定是跨库兼容问题,还是缺少正确的标题、列表检测方法,以下是相关内容,寻求解决方案:

输入Docx内容

C++ <-- header
C++ is a general purpose, high-level programming language developed by Sun Microsystems. 
The C++ programming language was developed by a small team of engineers, 
known as the Green Team, who initiated the language in 1991.
- Chapter 1
- Chapter 2
- Chapter 3

转换器代码

package lab2.converter;

import java.io.*;
import java.util.*;
import org.apache.poi.xwpf.usermodel.*;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;

public class HtmlConverter implements Converter{

    @Override
    public void convert() throws IOException {
        try {
            // 1. 用Apache POI读取Word文档
            FileInputStream fis = new FileInputStream("output/Find&Replace.docx");
            XWPFDocument document = new XWPFDocument(fis);

            // 2. 用jsoup转成HTML
            Document htmlDoc = new Document("");
            Element html = htmlDoc.appendElement("html");
            Element body = html.appendElement("body");

            List<XWPFParagraph> paragraphs = document.getParagraphs();
            for (XWPFParagraph paragraph : paragraphs) {
                Element p = body.appendElement("p");
                String text = paragraph.getText();

                // 检测标题和列表的判断条件
                // 未实现

                // 把文本添加到段落
                p.appendText(text);
            }

            // 3. 把HTML写入文件
            try (PrintWriter out = new PrintWriter("output/Find&Replace.html")) {
                out.println(htmlDoc.html());
            }

            System.out.println("转换完成。");
            document.close();

        } catch (Exception e) {
            System.err.println("Word转HTML出错:" + e.getMessage());
        }

    }
}

转换后的HTML结果

<html>
 <body>
  <p>Evaluation Warning: The document was created with Spire.Doc for JAVA.</p>
  <p>Evaluation Warning: The document was created with Spire.Doc for C++.</p>
  <p>C++</p>
  <p>C++ is a general purpose, high-level programming language developed by Sun Microsystems. The C++ programming language was developed by a small team of engineers, known as the Green Team, who initiated the language in 1991.</p>
  <p>Chapter 1</p>
  <p>Chapter 2</p>
  <p>Chapter 3</p>
 </body>
</html>

解决方案

1. 标题识别:通过样式或格式判断

Spire.Doc生成的标题会对应Word内置的标题样式(如Heading 1),可通过XWPFParagraph的样式信息或字体格式识别:

方法一:通过样式ID/名称判断

for (XWPFParagraph paragraph : paragraphs) {
    String text = paragraph.getText().trim();
    // 跳过空段落和Spire评估警告
    if (text.isEmpty() || text.startsWith("Evaluation Warning")) {
        continue;
    }

    Element contentElement;
    // 检查段落样式
    String styleId = paragraph.getStyleID();
    String styleName = paragraph.getStyle() != null ? paragraph.getStyle().getName() : "";
    
    // 匹配Word内置标题样式
    if ((styleId != null && styleId.startsWith("Heading")) || styleName.startsWith("Heading")) {
        // 根据标题级别生成h1-h6,这里以Heading 1为例对应h1
        contentElement = body.appendElement("h1");
    } else {
        contentElement = body.appendElement("p");
    }
    contentElement.appendText(text);
}

方法二:通过字体格式辅助判断

如果Spire生成的样式非标准Heading开头,可通过字体大小、加粗属性判断:

if (!paragraph.getRuns().isEmpty()) {
    XWPFRun firstRun = paragraph.getRuns().get(0);
    if (firstRun.isBold() && firstRun.getFontSize() >= 16) {
        contentElement = body.appendElement("h1");
    }
}

2. 列表识别:检查段落的列表格式

Apache POI可通过段落的getNumPr()属性判断是否为列表项,同时维护当前列表元素避免重复创建:

Element currentUl = null;
for (XWPFParagraph paragraph : paragraphs) {
    String text = paragraph.getText().trim();
    if (text.isEmpty() || text.startsWith("Evaluation Warning")) {
        continue;
    }

    // 判断是否为列表项
    boolean isListItem = paragraph.getCTP().getPPr() != null 
                        && paragraph.getCTP().getPPr().getNumPr() != null;
    
    if (isListItem) {
        if (currentUl == null) {
            currentUl = body.appendElement("ul");
        }
        // 移除文本开头的"- "标记
        currentUl.appendElement("li").appendText(text.replaceFirst("^- ", ""));
    } else {
        // 退出列表模式
        currentUl = null;
        // 处理标题或普通段落
        Element contentElement;
        String styleName = paragraph.getStyle() != null ? paragraph.getStyle().getName() : "";
        if (styleName.startsWith("Heading")) {
            contentElement = body.appendElement("h1");
        } else {
            contentElement = body.appendElement("p");
        }
        contentElement.appendText(text);
    }
}

3. 跨库兼容注意事项

Spire.Doc生成的Docx符合OOXML标准,Apache POI可正常解析,但需注意:

  • 免费版Spire会插入评估警告段落,需过滤
  • 部分自定义样式的ID/名称可能与POI默认识别规则有差异,必要时结合字体、段落格式辅助判断

内容的提问来源于stack exchange,提问作者apo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 09:15:31