You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用XSLT 3.0实现文献编号的交叉链接转换?

解决XSLT 3.0处理文献多编号与范围编号交叉链接的问题

当前使用XSLT 3.0为XML中文献编号添加交叉链接时,单个编号(如[1])转换正常,但多编号(如[6, 9, 12])和范围编号(如[8, 11-14, 17])无法按预期生成交叉链接,需调整XSLT代码实现需求。

输入XML

<?xml version="1.0"?>
<book id="bk1">
<p>The heterogeneity of patients, various clinical manifestations and the dynamics of CS development cause problems **[1]** with identifying its unified definition. However, CS can be usually diagnosed on the basis of clinical criteria which are easy to assess without the need for advanced hemodynamic monitoring **[6, 9, 12]**. Increasing knowledge about **[8, 11-14, 17]** patient characteristics and better understanding of the CS pathophysiology.</p>
</book>

预期XML

<?xml version="1.0"?>
<book id="bk1">
<p>The heterogeneity of patients, various clinical manifestations and the dynamics of CS development cause problems **[<a href="#bib1">1</a>]** with identifying its unified definition. However, CS can be usually diagnosed on the basis of clinical criteria which are easy to assess without the need for advanced hemodynamic monitoring **[<a href="#bib6">6</a>,  <a href="#bib9">9</a>, <a href="#bib12">12</a>]**. Increasing knowledge about **[<a href="#bib8">8</a>, <a href="#bib11">11</a><a href="#bib12"></a><a href="#bib13"></a>-<a href="#bib14">14</a>, <a href="#bib17">17</a>]** patient characteristics and better understanding of the CS pathophysiology.</p>
</book>

当前XSLT代码

<?xml version="1.0" encoding="utf-8"?>
 <xsl:stylesheet version="3.0"
 xmlns:xsl="http://www.w3.org/1999/XSL/Transform"
 xmlns:xs="http://www.w3.org/2001/XMLSchema"
 xmlns:fn="http://www.w3.org/2005/xpath-functions"
 exclude-result-prefixes="#all">
   <xsl:output method="xhtml" />
   <xsl:template match="/|node()|*|@*">
     <xsl:copy>
       <xsl:apply-templates select="node()|*|@*" />
     </xsl:copy>
   </xsl:template>
 
   <xsl:param name="para-rgxp">
     <text-patterns>
       <bibnosingle>\[(\d+)\]</bibnosingle>
       <bibunnumber>(\p{Lu}[\p{L}-]+)( et al.?)\s+\(([0-9]{4})\)</bibunnumber>
       <bibnomultiline>\[(\d+),\s+(\d+)</bibnomultiline>
     </text-patterns>
   </xsl:param>
 
   <xsl:template match="div[not(preceding-sibling::*[matches(.,'^References')])]/text()" priority="10">
     <xsl:analyze-string select="." regex="{string-join($para-rgxp/text-patterns/*,'|')}">
       <xsl:matching-substring>
  <xsl:choose>
    <xsl:when test="matches(.,$para-rgxp/text-patterns/bibnosingle)" >
      <a href="#bib{regex-group(1)}">
        <xsl:value-of select="."/>     
      </a>
    </xsl:when>
        <xsl:when test="matches(.,$para-rgxp/text-patterns/bibunnumber)" >
      <a href="#bib{regex-group(2)}">
        <xsl:value-of select="."/>     
      </a>
        </xsl:when>
        <xsl:when test="matches(.,$para-rgxp/text-patterns/bibnomultiline)" >
      <a href="#bib{regex-group(3)}{regex-group(4)}">
        <xsl:value-of select="."/>     
      </a>
        </xsl:when>
    <xsl:otherwise>
      <xsl:value-of select="."/>
    </xsl:otherwise>
  </xsl:choose>
       </xsl:matching-substring>
       <xsl:non-matching-substring>
  <xsl:value-of select="."/>
       </xsl:non-matching-substring>
     </xsl:analyze-string>
   </xsl:template>
</xsl:stylesheet>

修改后的XSLT代码

<?xml version="1.0" encoding="utf-8"?>
<xsl:stylesheet version="3.0"
  xmlns:xsl="http://www.w3.org/1999/XSL/Transform"
  xmlns:xs="http://www.w3.org/2001/XMLSchema"
  exclude-result-prefixes="#all">

  <xsl:output method="xhtml" />
  
  <!-- 身份模板 -->
  <xsl:template match="/|node()|@*">
    <xsl:copy>
      <xsl:apply-templates select="node()|@*" />
    </xsl:copy>
  </xsl:template>

  <!-- 匹配p元素下的文本节点(对应输入XML的结构) -->
  <xsl:template match="p/text()" priority="10">
    <!-- 第一层:匹配整个文献引用块 [xxx] -->
    <xsl:analyze-string select="." regex="\[([^\]]+)\]">
      <xsl:matching-substring>
        <xsl:text>[</xsl:text>
        <!-- 第二层:拆分引用块内的内容,分别处理编号、范围、分隔符 -->
        <xsl:analyze-string select="regex-group(1)" regex="(\d+)(-\d+)?|([^,\d]+)">
          <xsl:matching-substring>
            <!-- 处理单个数字编号 -->
            <xsl:if test="matches(., '^\d+$')">
              <a href="#bib{.}">
                <xsl:value-of select="."/>
              </a>
            </xsl:if>
            <!-- 处理范围编号(如11-14) -->
            <xsl:if test="matches(., '^(\d+)-(\d+)$')">
              <xsl:variable name="start" as="xs:integer" select="xs:integer(regex-group(1))"/>
              <xsl:variable name="end" as="xs:integer" select="xs:integer(regex-group(2))"/>
              <!-- 生成起始编号链接 -->
              <a href="#bib{$start}">
                <xsl:value-of select="$start"/>
              </a>
              <!-- 为中间编号生成空链接 -->
              <xsl:for-each select="$start + 1 to $end - 1">
                <a href="#bib{.}"></a>
              </xsl:for-each>
              <!-- 生成分隔符和结束编号链接 -->
              <xsl:text>-</xsl:text>
              <a href="#bib{$end}">
                <xsl:value-of select="$end"/>
              </a>
            </xsl:if>
            <!-- 保留分隔符(逗号、空格等) -->
            <xsl:if test="matches(., '^[^,\d]+$')">
              <xsl:value-of select="."/>
            </xsl:if>
          </xsl:matching-substring>
        </xsl:analyze-string>
        <xsl:text>]</xsl:text>
      </xsl:matching-substring>
      <xsl:non-matching-substring>
        <xsl:value-of select="."/>
      </xsl:non-matching-substring>
    </xsl:analyze-string>
  </xsl:template>

</xsl:stylesheet>

关键修改点说明

  1. 匹配目标修正:原代码匹配div下的文本,但输入XML中引用位于p元素内,因此将匹配规则改为p/text(),确保命中目标文本。
  2. 双层正则处理逻辑:
    • 第一层正则匹配整个引用块\[([^\]]+)\],提取括号内的所有内容。
    • 第二层拆分引用内容,分别处理单个编号、范围编号和分隔符,覆盖所有引用场景。
  3. 范围编号特殊处理:对x-y格式的范围,解析起始和结束数字,生成起始编号链接,为中间数字添加空链接,最后生成结束编号链接,完全匹配预期输出格式。
  4. 移除无效规则:原代码中bibnomultiline规则匹配不完整且逻辑错误,替换为通用拆分逻辑,无需依赖零散的正则片段。

内容的提问来源于stack exchange,提问作者Balaji

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 02:31:03