XSLT带if条件的节点选择为何选中全部而非指定节点属性
问题根因
原代码的逻辑错误直接导致输出异常和内存溢出:
<xsl:if>的判断逻辑是全局存在性校验:只要所有rel里存在任意一个name为is_comparable_to的节点、且存在任意一个rel的from等于当前unit的id,条件就会成立,没有校验两个属性是否属于同一个rel节点,导致分支大量误触发。- 分支触发后执行的
<xsl:value-of select="$all/relations/rel/@to"/>会选中所有源文件中所有rel节点的to属性,既会把所有属性值拼接成错误的输出内容,又会触发全量节点遍历,数据量稍大就会占满内存。 - 原代码模板匹配规则写的是
node()|@*,会递归匹配文档中所有节点和属性,产生大量无意义的执行,额外增加内存开销。 - 原RDF输出语法不符合规范,对象属性的目标资源URI应该直接写在
rdf:resource属性中,不需要把ID值放在标签文本节点内。
修复方案
1. 预过滤降低遍历开销
在遍历unit节点的循环外,提前把所有符合name要求的rel节点筛选出来存为变量,避免每次循环都全量扫描所有rel节点,把重复计算的开销降到最低:
<xsl:variable name="all" select="collection('./?select=*.xml')"/> <!-- 提前筛选所有可比关系,后续逻辑仅遍历这部分节点 --> <xsl:variable name="all_rels_comparable" select="$all/relations/rel[@name = 'is_comparable_to']"/>
2. 替换错误的分支逻辑
删除原有存在性判断的<xsl:if>块,将模板匹配规则改为匹配根节点/避免递归遍历,直接在预过滤的关系节点中,匹配from属性等于当前unit id的节点,用<xsl:for-each>遍历(支持一个unit关联多个可比关系的场景),直接输出符合RDF规范的属性节点:
<xsl:for-each select="$all/sets/singleSet/unit"> <xsl:variable name="id_unit" select="./@id"/> <xsl:variable name="uri_unit" select="concat('https://blabla/', $id_unit)"/> <NamedIndividual rdf:about="{$uri_unit}"> <bla:identifier rdf:datatype="http://www.w3.org/2001/XMLSchema#string"> <xsl:value-of select="./@id"/> </bla:identifier> <!-- 遍历当前unit关联的所有可比关系,生成标准RDF资源声明 --> <xsl:for-each select="$all_rels_comparable[@from = $id_unit]"> <bla:is_comparable_to rdf:resource="blabla/is_comparable_to/{./@to}" /> </xsl:for-each> </NamedIndividual> </xsl:for-each>
完整修复后的XSLT代码
<xsl:stylesheet version="3.0" xmlns:bla="blabla" xmlns:owl="http://www.w3.org/2002/07/owl#" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:xml="http://www.w3.org/XML/1998/namespace" xmlns:xsd="http://www.w3.org/2001/XMLSchema#" xmlns:rdfs="http://www.w3.org/2000/01/rdf-schema#"> <xsl:template match="/"> <rdf:RDF> <xsl:variable name="all" select="collection('./?select=*.xml')"/> <xsl:variable name="all_rels_comparable" select="$all/relations/rel[@name = 'is_comparable_to']"/> <xsl:for-each select="$all/sets/singleSet/unit"> <xsl:variable name="id_unit" select="./@id"/> <xsl:variable name="uri_unit" select="concat('https://blabla/', $id_unit)"/> <NamedIndividual rdf:about="{$uri_unit}"> <bla:identifier rdf:datatype="http://www.w3.org/2001/XMLSchema#string"> <xsl:value-of select="./@id"/> </bla:identifier> <xsl:for-each select="$all_rels_comparable[@from = $id_unit]"> <bla:is_comparable_to rdf:resource="blabla/is_comparable_to/{./@to}" /> </xsl:for-each> </NamedIndividual> </xsl:for-each> </rdf:RDF> </xsl:template> </xsl:stylesheet>
优化效果
- 时间复杂度从原来的O(nm)(n为unit总数,m为rel总数)降到O(np + m)(p为可比关系总数,远小于m),从根源上避免全量无效遍历导致的内存溢出问题。
- XPath谓词同时校验同一个rel节点的
name和from属性,不会出现误匹配,输出内容完全符合预期。 - 采用标准RDF/XML语法输出资源引用,不需要额外解析文本节点内容,兼容性更好。
内容的提问来源于stack exchange,提问作者piaschwarz
相关产品推荐
相关产品推荐

