HTML转XML的XSLT转换中如何拆分元素文本与节点且不丢失标记
问题描述
现有HTML输入:
<p> <i>This is italic tag</i>. This text belongs to x element. The text from here should <sub>2</sub> belong to y element. </p>
使用以下XSLT转换后出现内容丢失:
<xsl:template match="p/text()[not(parent::i)]"> <xsl:variable name="tokens" select="tokenize(.,'\.')"/> <xsl:element name="x"> <xsl:value-of select="normalize-space($tokens[2])"/> </xsl:element> <xsl:element name="y"> <xsl:value-of select="normalize-space(string-join($tokens[position() > 2 and position() != last()], '.'))"/> </xsl:element> </xsl:template> <xsl:template match="sub"> <xsl:element name="SUBSCRIPT"> <xsl:apply-templates/> </xsl:element> </xsl:template> <xsl:template match="p"> <xsl:element name="PARA"> <xsl:apply-templates/> </xsl:element> </xsl:template> <xsl:template match="i"> <xsl:element name="EMPHASIS"> <xsl:apply-templates/> </xsl:element> </xsl:template>
当前错误输出:
<PARA> <EMPHASIS>This is italic tag</EMPHASIS> <X>This text belongs to x element</X> <Y/> <SUBSCRIPT>2</SUBSCRIPT> <X/> <Y/> </PARA>
期望输出:
<PARA> <EMPHASIS>This is italic tag</EMPHASIS> <X>This text belongs to x element</X> <Y>The text from here should <SUBSCRIPT>2</SUBSCRIPT> belong to y element</Y> </PARA>
需要调整XSLT模板,实现拆分<p>标签文本并保留<sub>子元素,避免内容丢失。
解决方案
问题根源是原模板仅处理<p>下的纯文本节点,未处理文本与<sub>混合的节点组,导致<sub>被孤立处理,周围文本丢失。调整思路是先对<p>内除<i>外的所有内容按句号分组,再分别包装为<X>和<Y>,同时保留子元素的转换逻辑。
调整后的XSLT代码:
<xsl:template match="p"> <PARA> <!-- 优先处理i元素 --> <xsl:apply-templates select="i"/> <!-- 获取p下除i外的所有节点(文本+sub) --> <xsl:variable name="content-after-i" select="node()[not(self::i)]"/> <!-- 定位第二个句号所在的节点,以此分割X/Y内容 --> <xsl:variable name="split-point" select="$content-after-i[contains(., '.')][2]"/> <!-- 生成X元素:包含split-point之前的所有内容 --> <X> <xsl:apply-templates select="$content-after-i[. << $split-point]"> <xsl:with-param name="normalize" select="true()"/> </xsl:apply-templates> </X> <!-- 生成Y元素:包含split-point之后的所有内容 --> <Y> <xsl:apply-templates select="$content-after-i[. >> $split-point]"> <xsl:with-param name="normalize" select="true()"/> </xsl:apply-templates> </Y> </PARA> </xsl:template> <xsl:template match="sub"> <SUBSCRIPT> <xsl:apply-templates/> </SUBSCRIPT> </xsl:template> <xsl:template match="i"> <EMPHASIS> <xsl:apply-templates/> </EMPHASIS> </xsl:template> <!-- 通用文本节点处理模板,支持空格归一化控制 --> <xsl:template match="text()"> <xsl:param name="normalize" select="false()"/> <xsl:choose> <xsl:when test="$normalize"> <xsl:value-of select="normalize-space(.)"/> </xsl:when> <xsl:otherwise> <xsl:value-of select="."/> </xsl:otherwise> </xsl:choose> </xsl:template>
关键改动说明:
- 重构
<p>模板逻辑:不再单独处理纯文本节点,而是提取<p>内除<i>外的全部节点集合,通过节点位置运算符<</>>按第二个句号分割内容组。 - 保留子元素转换:对分组后的内容直接应用模板,确保
<sub>能被原有模板正确转换为<SUBSCRIPT>并嵌入对应位置。 - 可控的空格归一化:新增文本节点模板,通过参数控制是否执行
normalize-space,避免无差别空格处理导致的内容变形。
内容的提问来源于stack exchange,提问作者Bakhovuddin Latifjanov
相关产品推荐
相关产品推荐

