如何用XSLT 3按特定正则匹配拆分混合内容节点?
问题:按全大写单词拆分混合内容段落
输入XML
<stuff> <p>CAPITALWORD is part of <i>mixed</i> content.</p> <p>ANOTHER is <i>here</i> but it's not the only one. SOMEWORDS are <i>mixted up</i> in the same paragraph. SOMETIMES even <i>multiple times.</i></p> </stuff>
尝试的XSLT代码
<xsl:output method="xml" indent="true"></xsl:output> <xsl:mode on-no-match="shallow-copy"/> <xsl:template match="p"> <xsl:for-each-group select="node()" group-starting-with="text()[matches(., '[A-Z]{2,}')]"> <xsl:element name="p" > <xsl:apply-templates select="current-group()"/> </xsl:element> </xsl:for-each-group> </xsl:template>
当前错误输出
<stuff> <p>CAPITALWORD is part of <i>mixed</i> content.</p> <p>ANOTHER is <i>here</i> </p> <p> but it's not the only one. SOMEWORDS are <i>mixed up</i> in the <i>same</i> </p> <p> paragraph. SOMETIMES even <i>multiple times.</i> </p> </stuff>
期望输出
<stuff> <p>CAPITALWORD is part of <i>mixed</i> content. </p> <p>ANOTHER is <i>here</i> but it's not the only one. </p> <p>SOMEWORDS are <i>mixed up</i> in the <i>same</i> paragraph. </p> <p>SOMETIMES even <i>multiple times.</i></p> </stuff>
解决方案
原代码的问题在于,group-starting-with仅匹配整个文本节点是否包含大写单词,但大写单词往往只是文本节点的一部分,导致分组逻辑提前断开,把同一语义段的内容拆分到多个<p>中。
以下是修正后的XSLT 3.0代码,核心思路是先拆分文本节点,将每个全大写单词(≥2个大写字母)转为独立节点,再按这些节点分组生成新段落:
<xsl:stylesheet version="3.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"> <xsl:output method="xml" indent="true"/> <xsl:mode on-no-match="shallow-copy"/> <xsl:template match="p"> <!-- 先处理所有节点,拆分文本中的大写单词为独立节点 --> <xsl:variable name="processed-nodes"> <xsl:apply-templates select="node()" mode="split-uppercase"/> </xsl:variable> <!-- 按大写单词开头的文本节点分组,生成新段落 --> <xsl:for-each-group select="$processed-nodes/node()" group-starting-with="text()[matches(., '^[A-Z]{2,}')]"> <p> <xsl:apply-templates select="current-group()"/> </p> </xsl:for-each-group> </xsl:template> <!-- 拆分文本节点:将每个全大写单词转为独立文本节点 --> <xsl:template match="text()" mode="split-uppercase"> <xsl:analyze-string select="." regex="([A-Z]{2,})" flags="s"> <xsl:non-matching-substring> <xsl:if test="normalize-space(.)"> <xsl:value-of select="."/> </xsl:if> </xsl:non-matching-substring> <xsl:matching-substring> <xsl:value-of select="."/> </xsl:matching-substring> </xsl:analyze-string> </xsl:template> <!-- 保留元素节点结构,递归处理内部内容 --> <xsl:template match="*" mode="split-uppercase"> <xsl:copy> <xsl:apply-templates select="@*, node()" mode="split-uppercase"/> </xsl:copy> </xsl:template> </xsl:stylesheet>
代码说明
- 文本拆分:通过
xsl:analyze-string将原文本拆分为非大写单词片段和独立的大写单词节点,确保每个全大写单词都成为一个可被分组识别的独立节点。 - 分组生成段落:对处理后的所有节点,以大写单词开头的文本节点为分组起点,每个分组对应一个新的
<p>元素,完整保留原有的混合内容(如<i>元素)结构。
内容的提问来源于stack exchange,提问作者Tench
相关产品推荐
相关产品推荐

