You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用XSLT 3按特定正则匹配拆分混合内容节点?

问题:按全大写单词拆分混合内容段落

输入XML

<stuff>
    <p>CAPITALWORD is part of <i>mixed</i> content.</p>
    <p>ANOTHER is <i>here</i> but it's not the only one. SOMEWORDS are <i>mixted up</i> in the same
        paragraph. SOMETIMES even <i>multiple times.</i></p>
</stuff>

尝试的XSLT代码

<xsl:output method="xml" indent="true"></xsl:output>
<xsl:mode on-no-match="shallow-copy"/>
    
<xsl:template match="p">
  <xsl:for-each-group select="node()" group-starting-with="text()[matches(., '[A-Z]{2,}')]">
    <xsl:element name="p" >
      <xsl:apply-templates select="current-group()"/>
    </xsl:element>  
  </xsl:for-each-group>
</xsl:template>

当前错误输出

<stuff>
   <p>CAPITALWORD is part of <i>mixed</i> content.</p>
   <p>ANOTHER is <i>here</i>
   </p>
   <p> but it's not the only one. SOMEWORDS are <i>mixed up</i> in the <i>same</i>
   </p>
   <p>
        paragraph. SOMETIMES even <i>multiple times.</i>
   </p>
</stuff>

期望输出

<stuff>
    <p>CAPITALWORD is part of <i>mixed</i> content. </p>
    <p>ANOTHER is <i>here</i> but it's not the only one. </p>
    <p>SOMEWORDS are <i>mixed up</i> in the <i>same</i> paragraph. </p>
    <p>SOMETIMES even <i>multiple times.</i></p>
</stuff>

解决方案

原代码的问题在于,group-starting-with仅匹配整个文本节点是否包含大写单词,但大写单词往往只是文本节点的一部分,导致分组逻辑提前断开,把同一语义段的内容拆分到多个<p>中。

以下是修正后的XSLT 3.0代码,核心思路是先拆分文本节点,将每个全大写单词(≥2个大写字母)转为独立节点,再按这些节点分组生成新段落:

<xsl:stylesheet version="3.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform">
    <xsl:output method="xml" indent="true"/>
    <xsl:mode on-no-match="shallow-copy"/>

    <xsl:template match="p">
        <!-- 先处理所有节点,拆分文本中的大写单词为独立节点 -->
        <xsl:variable name="processed-nodes">
            <xsl:apply-templates select="node()" mode="split-uppercase"/>
        </xsl:variable>
        
        <!-- 按大写单词开头的文本节点分组,生成新段落 -->
        <xsl:for-each-group select="$processed-nodes/node()" 
                            group-starting-with="text()[matches(., '^[A-Z]{2,}')]">
            <p>
                <xsl:apply-templates select="current-group()"/>
            </p>
        </xsl:for-each-group>
    </xsl:template>

    <!-- 拆分文本节点:将每个全大写单词转为独立文本节点 -->
    <xsl:template match="text()" mode="split-uppercase">
        <xsl:analyze-string select="." regex="([A-Z]{2,})" flags="s">
            <xsl:non-matching-substring>
                <xsl:if test="normalize-space(.)">
                    <xsl:value-of select="."/>
                </xsl:if>
            </xsl:non-matching-substring>
            <xsl:matching-substring>
                <xsl:value-of select="."/>
            </xsl:matching-substring>
        </xsl:analyze-string>
    </xsl:template>

    <!-- 保留元素节点结构,递归处理内部内容 -->
    <xsl:template match="*" mode="split-uppercase">
        <xsl:copy>
            <xsl:apply-templates select="@*, node()" mode="split-uppercase"/>
        </xsl:copy>
    </xsl:template>
</xsl:stylesheet>

代码说明

  1. 文本拆分:通过xsl:analyze-string将原文本拆分为非大写单词片段和独立的大写单词节点,确保每个全大写单词都成为一个可被分组识别的独立节点。
  2. 分组生成段落:对处理后的所有节点,以大写单词开头的文本节点为分组起点,每个分组对应一个新的<p>元素,完整保留原有的混合内容(如<i>元素)结构。

内容的提问来源于stack exchange,提问作者Tench

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 02:32:47