You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用XSLT替换文本子串为元素并保留原有内嵌元素

问题

我有一份带文本的XML文档和一份单词列表XML文档,需要用XSLT实现三个目标:

  • 保留所有现有元素及属性(包括内嵌元素)
  • 识别外部单词列表中的所有单词
  • 将这些单词替换为带列表对应id引用的元素

目前我实现了部分功能,但无法整合出预期输出。

输入XML

<doc>
    <header>Document with example sentences</header>
    <text>
        <div type="sentence" n="1">They<note>buyers</note> bought an apple and a banana.</div>
        <div type="sentence" n="2">They<note>shop</note> only had a strawberry and an apple left.</div>
    </text>
</doc>

单词列表XML(liste.xml)

<list>
    <fruit id="001">
        <english>apple</english>
        <translations>Apfel, pomme</translations>
    </fruit>
    <fruit id="002">
        <english>banana</english>
        <translations>Banane, banane</translations>
    </fruit>
    <fruit id="003">
        <english>strawberry</english>
        <translations>Erdbeere, strawberry</translations>
    </fruit>
</list>

期望输出XML

<doc>
    <header>Document with example sentences</header>
    <text>
        <div type="sentence" n="1">They<note>buyers</note> bought <fruit ref="#001">apple</fruit> and a <fruit ref="#002">banana</fruit>.</div>
        <div type="sentence" n="2"> They<note>shop</note> only had <fruit ref="#003">strawberry</fruit> and <fruit ref="#001">apple</fruit>left.</div>
    </text>
</doc>

我尝试了两种方案:第一种能识别文本中的列表单词,第二种能将单词替换为元素,但无法同时实现两者并保留所有元素。

识别单词的代码

<xsl:template match="/ | @*|node()">
    <xsl:copy>
        <xsl:apply-templates select="@*|node()"/>
    </xsl:copy>
</xsl:template> 

<xsl:variable name="list" select="document('liste.xml')"/>

<xsl:template match="div">
    <xsl:variable name="text" select="."/>
    
    <xsl:copy>    
        
        <xsl:apply-templates select="@*|node()"/>
    
        <xsl:for-each select="$list/list/fruit">
            <xsl:variable name="english" select="english"/>
            <xsl:if test="contains($text,$english)">
                <xsl:element name="identified_fruit">
                    <xsl:value-of select="$english"/>
                </xsl:element>
            </xsl:if>
        </xsl:for-each>
    </xsl:copy>
    
</xsl:template>

替换单词为元素的代码

<xsl:template match="@* | node()">
    <xsl:copy>
        <xsl:apply-templates select="@*, node()"/>
    </xsl:copy>
</xsl:template>

<xsl:template match="div">
    <xsl:copy>
        <xsl:apply-templates select="@*"/>
        <xsl:apply-templates select="text()" mode="wrap">
            <xsl:with-param name="words" as="xs:string+" select="'banana', 'apple', 'strawberry'"/>
        </xsl:apply-templates>
    </xsl:copy>
</xsl:template>

<xsl:template match="text()" mode="wrap">
    <xsl:param name="words" as="xs:string+"/>
    <xsl:param name="wrapper-name" as="xs:string" select="'fruit'"/>
    <xsl:analyze-string select="." regex="{string-join($words, '|')}">
        <xsl:matching-substring>
            <xsl:element name="{$wrapper-name}">
                <xsl:value-of select="."/>
            </xsl:element>
        </xsl:matching-substring>
        <non-matching-substring>
            <xsl:value-of select="."/>
        </non-matching-substring>
    </xsl:analyze-string>
</xsl:template>

请问如何整合这两种方案,同时保留文档中的所有元素?


解决方案

核心思路是保留通用的身份转换模板(用于复制所有原有元素和属性),然后针对文本节点做处理——从外部单词列表动态生成正则表达式,匹配单词后替换为带ref属性的fruit元素,同时保留内嵌元素的结构。

整合后的完整XSLT代码如下:

<xsl:stylesheet version="2.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"
                xmlns:xs="http://www.w3.org/2001/XMLSchema">

    <!-- 通用身份转换模板:复制所有原有元素、属性和文本(除了被后续模板匹配的文本节点) -->
    <xsl:template match="@* | node()">
        <xsl:copy>
            <xsl:apply-templates select="@* | node()"/>
        </xsl:copy>
    </xsl:template>

    <!-- 加载外部单词列表 -->
    <xsl:variable name="fruit-list" select="document('liste.xml')/list/fruit"/>
    <!-- 从单词列表生成正则表达式:匹配完整单词,避免部分匹配(比如apple不匹配apples) -->
    <xsl:variable name="fruit-regex" select="concat('\b(', string-join($fruit-list/english, '|'), ')\b')"/>

    <!-- 处理所有文本节点,替换匹配的水果单词 -->
    <xsl:template match="text()">
        <xsl:analyze-string select="." regex="{$fruit-regex}">
            <xsl:matching-substring>
                <!-- 根据匹配到的单词,找到对应的fruit节点,获取id -->
                <xsl:variable name="matched-word" select="."/>
                <xsl:variable name="target-fruit" select="$fruit-list[english = $matched-word]"/>
                <fruit ref="#{$target-fruit/@id}">
                    <xsl:value-of select="$matched-word"/>
                </fruit>
            </xsl:matching-substring>
            <non-matching-substring>
                <!-- 保留不匹配的文本内容 -->
                <xsl:value-of select="."/>
            </non-matching-substring>
        </xsl:analyze-string>
    </xsl:template>

</xsl:stylesheet>

关键说明

  1. 身份转换模板:确保所有原有元素(包括<note>这类内嵌元素)、属性都被完整复制,不会丢失结构。
  2. 动态正则生成:从外部单词列表提取所有英文单词,拼接成正则表达式,用\b确保匹配完整单词(若需要支持复数可调整正则规则)。
  3. 文本节点处理:用<xsl:analyze-string>拆分文本,匹配到单词时,通过单词找到对应列表项的id,生成带ref属性的<fruit>元素;不匹配的文本直接保留。
  4. 外部文档加载:通过document('liste.xml')加载单词列表,确保动态获取最新的单词和对应id。

这样就能同时满足三个需求:保留原有结构、识别外部单词、替换为带引用的元素。

内容的提问来源于stack exchange,提问作者RaBa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 19:44:55