You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用XSLT按ID分组,将相邻同元素的所有部分合并为单一内容

针对你提出的用XSLT按ID分组并合并相邻相同元素内容的需求,我整理了具体的实现思路和示例代码,一起来看看吧:

核心思路

我们需要利用XSLT的分组功能,先通过ID的前缀识别元素所属的分组,再把相邻的、同元素类型、同分组的元素合并成一个,将它们的子内容整合到一起。这里推荐用XSLT 2.0及以上版本的<xsl:for-each-group>特性,它能轻松处理相邻分组的场景。

输入XML示例

先把你提供的输入(补充了一个相邻同组的figure作为测试场景)整理如下:

<?xml version="1.0"?>
<he>
  <p>...some text 111.</p>
  <figure id="c01-fig-001">
    <section id="c01-sec-1418">
      <title type="main">A Journey</title>
      <section id="c01-sec-1218">
        <p>My para.</p>
      </section>
    </section>
    <caption>My Fig 1.</caption>
  </figure>
  <figure id="c01-fig-002" role="noart">
    <section id="c01-sec-1518">
      <title type="main">Another Journey</title>
      <p>Another para.</p>
    </section>
    <caption>My Fig 2.</caption>
  </figure>
  <figure id="c01-fig-003">
    <caption>My Fig 3.</caption>
  </figure>
  <p>...some text 222.</p>
</he>
完整XSLT实现
<?xml version="1.0" encoding="UTF-8"?>
<xsl:stylesheet version="2.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform">
  <xsl:output method="xml" indent="yes"/>

  <!-- 匹配根节点,处理所有子元素 -->
  <xsl:template match="/he">
    <he>
      <!-- 按「元素名+ID前缀」进行相邻分组 -->
      <xsl:for-each-group select="*" group-adjacent="concat(local-name(), '|', replace(@id, '^([^-]+-[^-]+)-.*$', '$1'))">
        <!-- 处理带ID的分组元素 -->
        <xsl:if test="current-group()[1]/@id">
          <xsl:element name="{local-name(current-group()[1])}">
            <!-- 保留第一个元素的所有属性 -->
            <xsl:copy-of select="current-group()[1]/@*"/>
            <!-- 合并组内所有元素的子内容 -->
            <xsl:copy-of select="current-group()/*"/>
          </xsl:element>
        </xsl:if>
        <!-- 不带ID的元素(比如p)直接原样输出 -->
        <xsl:if test="not(current-group()[1]/@id)">
          <xsl:copy-of select="current-group()"/>
        </xsl:if>
      </xsl:for-each-group>
    </he>
  </xsl:template>
</xsl:stylesheet>
代码关键说明
  1. 分组规则:group-adjacent="concat(local-name(), '|', replace(@id, '^([^-]+-[^-]+)-.*$', '$1'))"
    • local-name()获取元素的标签名(比如figure、section),确保只有相同类型的元素才会被分组。
    • replace(...)用正则提取ID的前缀:比如c01-fig-001会被提取为c01-fig,保证同分组的元素被归到一组。
  2. 合并逻辑:
    • 为每个分组创建一个新的同类型元素,保留分组内第一个元素的所有属性(避免属性冲突)。
    • 把分组内所有元素的子节点全部复制进去,实现内容的合并。
  3. 特殊元素处理:不带ID的元素(比如示例中的<p>)直接原样输出,不参与分组合并。
输出结果示例

运行上述XSLT后,相邻的三个c01-fig分组的figure会被合并成一个,输出如下:

<?xml version="1.0" encoding="UTF-8"?>
<he>
   <p>...some text 111.</p>
   <figure id="c01-fig-001">
      <section id="c01-sec-1418">
         <title type="main">A Journey</title>
         <section id="c01-sec-1218">
            <p>My para.</p>
         </section>
      </section>
      <caption>My Fig 1.</caption>
      <section id="c01-sec-1518">
         <title type="main">Another Journey</title>
         <p>Another para.</p>
      </section>
      <caption>My Fig 2.</caption>
      <caption>My Fig 3.</caption>
   </figure>
   <p>...some text 222.</p>
</he>
额外提示
  • 如果你的XSLT处理器只支持1.0版本,需要用key结合递归模板来实现相邻分组,不过现在主流处理器(比如Saxon)都支持2.0+,优先用上面的方案更简洁。
  • 如果ID格式有变化(比如用下划线分隔),只需要修改正则表达式即可,比如把replace里的-换成_。

内容的提问来源于stack exchange,提问作者sspsujit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:28:03