You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用XSLT提取DITA文件中两个处理指令之间的所有文本(含元素内文本)

XSLT 提取处理指令对之间内容的解决方案

核心逻辑

要避免跨组匹配,需要为每一个<?PI start?>处理指令,定位其同父节点下最近的后续<?PI end?>处理指令,仅提取两个节点之间的所有兄弟节点,即可保证范围准确。

方案1:适配Saxon-HE的XSLT 2.0方案(推荐)

Saxon-HE 9.x及以上版本原生支持XSLT 2.0,使用intersect运算符可以快速定位两个节点之间的内容:

<?xml version="1.0" encoding="UTF-8"?>
<xsl:stylesheet version="2.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform">
    <xsl:output method="xml" indent="yes"/>
    
    <!-- 身份复制模板,保留原XML结构,仅提取内容可删除 -->
    <xsl:template match="@*|node()">
        <xsl:copy>
            <xsl:apply-templates select="@*|node()"/>
        </xsl:copy>
    </xsl:template>
    
    <!-- 匹配start类型的PI -->
    <xsl:template match="processing-instruction('PI')[contains(., 'start')]">
        <!-- 定位同组对应的end PI -->
        <xsl:variable name="end-pi" select="following-sibling::processing-instruction('PI')[contains(., 'end')][1]"/>
        <!-- 取两个PI之间的所有节点 -->
        <xsl:variable name="content-between" select="following-sibling::node() intersect $end-pi/preceding-sibling::node()"/>
        
        <!-- 输出提取的文本,要保留元素结构替换为<xsl:copy-of select="$content-between"/> -->
        <xsl:value-of select="$content-between" separator=""/>
        
        <!-- 保留原PI标签,不需要可删除 -->
        <xsl:copy/>
    </xsl:template>
    
    <!-- 匹配end类型的PI,不需要保留可设为空模板 -->
    <xsl:template match="processing-instruction('PI')[contains(., 'end')]">
        <xsl:copy/>
    </xsl:template>
</xsl:stylesheet>

方案2:兼容XSLT 1.0的递归方案

如果必须使用XSLT 1.0,可以用递归遍历节点实现终止判断:

<?xml version="1.0" encoding="UTF-8"?>
<xsl:stylesheet version="1.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform">
    <xsl:output method="xml" indent="yes"/>
    
    <xsl:template match="@*|node()">
        <xsl:copy>
            <xsl:apply-templates select="@*|node()"/>
        </xsl:copy>
    </xsl:template>
    
    <xsl:template match="processing-instruction('PI')[contains(., 'start')]">
        <xsl:variable name="end-pi" select="following-sibling::processing-instruction('PI')[contains(., 'end')][1]"/>
        <!-- 调用递归模板提取中间内容 -->
        <xsl:apply-templates select="following-sibling::node()[1]" mode="extract">
            <xsl:with-param name="end-pi" select="$end-pi"/>
        </xsl:apply-templates>
        <xsl:copy/>
    </xsl:template>
    
    <!-- 递归提取模板,遇到end PI自动终止 -->
    <xsl:template match="node()" mode="extract">
        <xsl:param name="end-pi"/>
        <xsl:if test="generate-id(.) != generate-id($end-pi)">
            <!-- 输出当前节点文本,要保留元素结构替换为<xsl:copy-of select="."/> -->
            <xsl:value-of select="."/>
            <xsl:apply-templates select="following-sibling::node()[1]" mode="extract">
                <xsl:with-param name="end-pi" select="$end-pi"/>
            </xsl:apply-templates>
        </xsl:if>
    </xsl:template>
    
    <xsl:template match="processing-instruction('PI')[contains(., 'end')]">
        <xsl:copy/>
    </xsl:template>
</xsl:stylesheet>

效果验证

针对你的输入示例,两种方案都只会提取每对PI之间的内容,不会匹配到text03这类不属于当前PI对的内容,输出结果符合需求。

内容的提问来源于stack exchange,提问作者Plutto313

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 04:45:03