XSL开发需求:将XML流式文本中的列表转换为标准列表元素
Got it, let's figure out how to fix this grouping issue for your XML lists. The problem you're facing is super common when dealing with free-flowing text that has list items scattered as individual nodes—your existing XSL isn't bundling those consecutive list items into proper <ol> or <ul> parent elements, right? Let's break down the solution step by step.
Core Problem Breakdown
Your XML has list items as separate nodes (like <p> or <text>) with prefixes like 1., 2), -, or *, but they're not wrapped in parent list elements. We need to:
- Identify which nodes are list items (and whether they're ordered/unordered)
- Group consecutive list items of the same type
- Wrap each group in the correct
<ol>or<ul>element, and strip the prefix from each list item to make<li>elements - Leave non-list content untouched
Example Input & Expected Output
Let's use a sample to ground this:
Input XML
<content> <p>Here's some text before the lists:</p> <p>1. First numbered item</p> <p>2. Second numbered item</p> <p>3) Third numbered item (using parentheses)</p> <p>Random text between list blocks.</p> <p>- First bulleted item</p> <p>* Second bulleted item</p> <p>Final text after all lists.</p> </content>
Expected Output XML
<content> <p>Here's some text before the lists:</p> <ol> <li>First numbered item</li> <li>Second numbered item</li> <li>Third numbered item (using parentheses)</li> </ol> <p>Random text between list blocks.</p> <ul> <li>First bulleted item</li> <li>Second bulleted item</li> </ul> <p>Final text after all lists.</p> </content>
XSLT Solution (2.0+)
We'll use XSLT's xsl:for-each-group with group-adjacent to cluster consecutive list items, plus regex to identify list types and strip prefixes.
<?xml version="1.0" encoding="UTF-8"?> <xsl:stylesheet version="2.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"> <xsl:output method="xml" indent="yes"/> <!-- Match your root content element --> <xsl:template match="content"> <xsl:copy> <!-- Group adjacent nodes by their list type (ordered/unordered/non-list) --> <xsl:for-each-group select="*" group-adjacent="determine-list-type(.)"> <xsl:choose> <!-- Wrap ordered list groups in <ol> --> <xsl:when test="current-grouping-key() = 'ordered'"> <ol> <xsl:apply-templates select="current-group()" mode="ordered-item"/> </ol> </xsl:when> <!-- Wrap unordered list groups in <ul> --> <xsl:when test="current-grouping-key() = 'unordered'"> <ul> <xsl:apply-templates select="current-group()" mode="unordered-item"/> </ul> </xsl:when> <!-- Leave non-list content as-is --> <xsl:otherwise> <xsl:copy-of select="current-group()"/> </xsl:otherwise> </xsl:choose> </xsl:for-each-group> </xsl:copy> </xsl:template> <!-- Helper function to classify each node --> <xsl:function name="determine-list-type"> <xsl:param name="node"/> <xsl:variable name="clean-text" select="normalize-space($node/text())"/> <xsl:choose> <!-- Match ordered list prefixes: digits + . or ) --> <xsl:when test="matches($clean-text, '^\d+[\.)] ')">ordered</xsl:when> <!-- Match unordered list prefixes: - or * + space --> <xsl:when test="matches($clean-text, '^[\-\*] ')">unordered</xsl:when> <!-- Not a list item --> <xsl:otherwise>non-list</xsl:otherwise> </xsl:choose> </xsl:function> <!-- Strip ordered list prefixes and create <li> --> <xsl:template match="*" mode="ordered-item"> <li> <xsl:value-of select="replace(normalize-space(text()), '^\d+[\.)] ', '')"/> </li> </xsl:template> <!-- Strip unordered list prefixes and create <li> --> <xsl:template match="*" mode="unordered-item"> <li> <xsl:value-of select="replace(normalize-space(text()), '^[\-\*] ', '')"/> </li> </xsl:template> <!-- Default template to copy other nodes/attributes unchanged --> <xsl:template match="@*|node()"> <xsl:copy> <xsl:apply-templates select="@*|node()"/> </xsl:copy> </xsl:template> </xsl:stylesheet>
Key Details to Note
- Grouping Logic:
group-adjacentensures we only bundle consecutive list items—if there's non-list text between two list blocks, they'll become separate<ol>/<ul>elements (which is what you want for correct structure). - Regex Flexibility: The regex patterns can be extended if you need to support more list formats (like uppercase letters
A., Roman numeralsi., etc.). Just update thematches()conditions in thedetermine-list-typefunction. - Node Agnosticism: This code assumes list items are in
<p>nodes, but you can adjust the template matches (e.g.,match="text"instead ofmatch="*") if your XML uses different element names for text nodes.
内容的提问来源于stack exchange,提问作者Gunilla

