Apache Tika转换复杂Word为XHTML:保留列表样式的技术问询
I’ve run into exactly this frustrating issue with Apache Tika before—default conversions turn Word lists into styled <p> tags instead of proper <ul>/<ol>/<li> structures, which breaks semantic XHTML and makes styling a nightmare. Here are the most reliable fixes I’ve tested for complex docs with nested lists, tables, and images:
1. Customize Tika’s Style-to-HTML Mapping
Tika relies on a style mapping system to translate Word paragraph styles to HTML elements. You can override this behavior to detect list-specific styles and output proper list tags instead of generic paragraphs.
Here’s a simplified example of extending XHTMLContentHandler to handle list items:
import org.apache.tika.sax.XHTMLContentHandler; import org.xml.sax.Attributes; import org.xml.sax.helpers.AttributesImpl; public class ListAwareXHTMLHandler extends XHTMLContentHandler { private boolean inList = false; private int currentListLevel = 0; @Override public void startElement(String uri, String localName, String qName, Attributes atts) throws SAXException { // Check if we're dealing with a list paragraph (match your "list_Paragraph" class) if ("p".equals(qName) && "list_Paragraph".equals(atts.getValue("class"))) { // Initialize a list if we aren't already in one if (!inList) { super.startElement(uri, "ul", "ul", new AttributesImpl()); inList = true; currentListLevel++; } // Replace the <p> tag with <li> super.startElement(uri, "li", "li", atts); return; } // Handle nested lists by checking indentation or style variations (e.g., "list_Paragraph_2") if ("p".equals(qName) && "list_Paragraph_2".equals(atts.getValue("class"))) { if (currentListLevel == 1) { super.startElement(uri, "ul", "ul", new AttributesImpl()); currentListLevel++; } super.startElement(uri, "li", "li", atts); return; } super.startElement(uri, localName, qName, atts); } @Override public void endElement(String uri, String localName, String qName) throws SAXException { if ("p".equals(qName) && inList) { // End the <li> instead of <p> super.endElement(uri, "li", "li"); // Add logic to close lists when a non-list paragraph is encountered // (you'll need to track the next element's style to trigger this) return; } super.endElement(uri, localName, qName); } }
Tweak this to distinguish between ordered/unordered lists by checking Word’s internal list properties instead of just class names.
2. Post-Process the Generated XHTML
If you don’t want to dig into Tika’s internals, a quicker workaround is to parse the raw XHTML output and rewrite list elements using a library like Jsoup.
Example with Jsoup:
import org.jsoup.Jsoup; import org.jsoup.nodes.Document; import org.jsoup.nodes.Element; import org.jsoup.select.Elements; public class XhtmlListFixer { public static String fixListStructure(String rawXhtml) { Document doc = Jsoup.parse(rawXhtml); Elements listParagraphs = doc.select("p.list_Paragraph"); if (!listParagraphs.isEmpty()) { // Wrap consecutive list items in a <ul> Element parentUl = doc.createElement("ul"); listParagraphs.first().before(parentUl); for (Element p : listParagraphs) { Element li = doc.createElement("li"); // Preserve all content (including nested text/inline styles) li.html(p.html()); parentUl.appendChild(li); p.remove(); } } // Add logic here for nested lists by checking indentation or sub-classes return doc.html(); } }
This is simpler but less robust for deeply nested lists—you’ll need to add checks for indentation levels or style variations to create nested <ul>/<ol> structures.
3. Enable Tika’s Built-in List Parsing (Version-Dependent)
Some newer versions of Tika’s OOXML parser include a configuration flag to enable proper list extraction. Check if your version supports the parseLists parameter:
import org.apache.tika.parser.microsoft.ooxml.OOXMLParser; import org.apache.tika.parser.Parser; Parser ooxmlParser = new OOXMLParser(); // Enable native list parsing to output semantic <ul>/<ol>/<li> tags ((OOXMLParser) ooxmlParser).setParseLists(true);
This tells Tika to recognize Word’s native list structures instead of treating list items as regular styled paragraphs.
Pro Tips for Complex Docs
- Nested Lists: Track list levels using Word’s internal indentation values or hierarchical style names (e.g., "List Level 1", "List Level 2") to build proper nested HTML lists.
- Tables/Images: Tika usually preserves these elements well by default, but if styles go missing, extend the style mapping approach to retain table borders, image captions, and other formatting.
内容的提问来源于stack exchange,提问作者Sgotenks

