You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Apache Tika转换复杂Word为XHTML:保留列表样式的技术问询

Fixing Apache Tika's Incorrect List Output When Converting Complex Word Docs to XHTML

I’ve run into exactly this frustrating issue with Apache Tika before—default conversions turn Word lists into styled <p> tags instead of proper <ul>/<ol>/<li> structures, which breaks semantic XHTML and makes styling a nightmare. Here are the most reliable fixes I’ve tested for complex docs with nested lists, tables, and images:

1. Customize Tika’s Style-to-HTML Mapping

Tika relies on a style mapping system to translate Word paragraph styles to HTML elements. You can override this behavior to detect list-specific styles and output proper list tags instead of generic paragraphs.

Here’s a simplified example of extending XHTMLContentHandler to handle list items:

import org.apache.tika.sax.XHTMLContentHandler;
import org.xml.sax.Attributes;
import org.xml.sax.helpers.AttributesImpl;

public class ListAwareXHTMLHandler extends XHTMLContentHandler {
    private boolean inList = false;
    private int currentListLevel = 0;

    @Override
    public void startElement(String uri, String localName, String qName, Attributes atts) throws SAXException {
        // Check if we're dealing with a list paragraph (match your "list_Paragraph" class)
        if ("p".equals(qName) && "list_Paragraph".equals(atts.getValue("class"))) {
            // Initialize a list if we aren't already in one
            if (!inList) {
                super.startElement(uri, "ul", "ul", new AttributesImpl());
                inList = true;
                currentListLevel++;
            }
            // Replace the <p> tag with <li>
            super.startElement(uri, "li", "li", atts);
            return;
        }
        // Handle nested lists by checking indentation or style variations (e.g., "list_Paragraph_2")
        if ("p".equals(qName) && "list_Paragraph_2".equals(atts.getValue("class"))) {
            if (currentListLevel == 1) {
                super.startElement(uri, "ul", "ul", new AttributesImpl());
                currentListLevel++;
            }
            super.startElement(uri, "li", "li", atts);
            return;
        }
        super.startElement(uri, localName, qName, atts);
    }

    @Override
    public void endElement(String uri, String localName, String qName) throws SAXException {
        if ("p".equals(qName) && inList) {
            // End the <li> instead of <p>
            super.endElement(uri, "li", "li");
            // Add logic to close lists when a non-list paragraph is encountered
            // (you'll need to track the next element's style to trigger this)
            return;
        }
        super.endElement(uri, localName, qName);
    }
}

Tweak this to distinguish between ordered/unordered lists by checking Word’s internal list properties instead of just class names.

2. Post-Process the Generated XHTML

If you don’t want to dig into Tika’s internals, a quicker workaround is to parse the raw XHTML output and rewrite list elements using a library like Jsoup.

Example with Jsoup:

import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;
import org.jsoup.select.Elements;

public class XhtmlListFixer {
    public static String fixListStructure(String rawXhtml) {
        Document doc = Jsoup.parse(rawXhtml);
        Elements listParagraphs = doc.select("p.list_Paragraph");
        
        if (!listParagraphs.isEmpty()) {
            // Wrap consecutive list items in a <ul>
            Element parentUl = doc.createElement("ul");
            listParagraphs.first().before(parentUl);
            
            for (Element p : listParagraphs) {
                Element li = doc.createElement("li");
                // Preserve all content (including nested text/inline styles)
                li.html(p.html());
                parentUl.appendChild(li);
                p.remove();
            }
        }
        
        // Add logic here for nested lists by checking indentation or sub-classes
        return doc.html();
    }
}

This is simpler but less robust for deeply nested lists—you’ll need to add checks for indentation levels or style variations to create nested <ul>/<ol> structures.

3. Enable Tika’s Built-in List Parsing (Version-Dependent)

Some newer versions of Tika’s OOXML parser include a configuration flag to enable proper list extraction. Check if your version supports the parseLists parameter:

import org.apache.tika.parser.microsoft.ooxml.OOXMLParser;
import org.apache.tika.parser.Parser;

Parser ooxmlParser = new OOXMLParser();
// Enable native list parsing to output semantic <ul>/<ol>/<li> tags
((OOXMLParser) ooxmlParser).setParseLists(true);

This tells Tika to recognize Word’s native list structures instead of treating list items as regular styled paragraphs.

Pro Tips for Complex Docs

  • Nested Lists: Track list levels using Word’s internal indentation values or hierarchical style names (e.g., "List Level 1", "List Level 2") to build proper nested HTML lists.
  • Tables/Images: Tika usually preserves these elements well by default, but if styles go missing, extend the style mapping approach to retain table borders, image captions, and other formatting.

内容的提问来源于stack exchange,提问作者Sgotenks

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:50:22