You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

XSLT字符串转节点及disable-output-escaping使用合理性咨询

问题:Access导出XML中嵌套转义标签的处理方案优化咨询

我正在将Access数据库导出为XML,之后需转换XML以准备数据供Framemaker创建出版物。过程中需在Access输出文本内创建交叉引用,插入<idref>some text</idref>后,导出XML时该标签会变成&amp;amp;lt;some text&amp;amp;gt;,导致HTML解析器无法将其转为目标节点。我使用Saxon EE 9.8.3及David Carlisle的htmlparse工具,现有方案采用disable-output-escaping,想咨询这是否为最优解决办法。

输入XML示例

<?xml version="1.0" encoding="UTF-8"?>
<dataroot xmlns:od="urn:schemas-microsoft-com:officedata" generated="2023-09-29T08:29:47">
<TEQuery>
<IntID>PR090F</IntID>
<TEName>Exempt Lease From Taxable Owner</TEName>
<Description>
&lt;div&gt;&lt;font face=&quot;Times New Roman&quot; color=black&gt;&amp;nbsp;Leased &amp;lt;idref&amp;gt;PR001F&amp;lt;/idref&amp;gt; properties that qualify for this exemption are reported under one of the following expenditures: &lt;/font&gt;&lt;/div&gt;

&lt;ul&gt;
 &lt;ul&gt;
  &lt;ul&gt;
   &lt;ul&gt;
    &lt;ul&gt;
     &lt;ul&gt;
      &lt;ul&gt;
       &lt;li&gt;&lt;font face=&quot;Times New Roman&quot; color=black&gt;&amp;lt;idref&amp;gt;PR001F&amp;lt;/idref&amp;gt;, &lt;/font&gt;&lt;/li&gt;
       &lt;li&gt;&lt;font face=&quot;Times New Roman&quot; color=black&gt;PR007F, &lt;/font&gt;&lt;/li&gt;
       &lt;li&gt;&lt;font face=&quot;Times New Roman&quot; color=black&gt;PR079F, &lt;/font&gt;&lt;/li&gt;
       &lt;li&gt;&lt;font face=&quot;Times New Roman&quot; color=black&gt;PR083F, &lt;/font&gt;&lt;/li&gt;
       &lt;li&gt;&lt;font face=&quot;Times New Roman&quot; color=black&gt;PR085F, &lt;/font&gt;&lt;/li&gt;
       &lt;li&gt;&lt;font face=&quot;Times New Roman&quot; color=black&gt;PR086F, &lt;/font&gt;&lt;/li&gt;
       &lt;li&gt;&lt;font face=&quot;Times New Roman&quot; color=black&gt;PR087F, .&lt;/font&gt;&lt;/li&gt;
      &lt;/ul&gt;
     &lt;/ul&gt;
    &lt;/ul&gt;
   &lt;/ul&gt;
  &lt;/ul&gt;
 &lt;/ul&gt;
&lt;/ul&gt;</Description>
<TaxSort>2</TaxSort>
</TEQuery>
</dataroot>

期望输出XML

<dataroot xmlns:od="urn:schemas-microsoft-com:officedata"
          generated="2023-09-26T10:37:15">

   <TaxExpenditure id="PR090F" TAXSORT="2">Exempt Lease From Taxable Owner
      <Description>
Leased <idref>PR001F</idref> properties that qualify for this exemption are reported under one of the following expenditures:
<unorderedList>
            <listitem><idref>PR001F</idref>, </listitem>
            <listitem>PR007F, </listitem>
            <listitem>PR079F, </listitem>
            <listitem>PR083F, </listitem>
            <listitem>PR085F, </listitem>
            <listitem>PR086F, </listitem>
            <listitem>PR087F, </listitem>
         </unorderedList>
   </TaxExpenditure>
</dataroot>

*注:使用的是支持XSLT 2.0的David Carlisle的htmlparse工具。

当前使用的XSLT代码

<xsl:stylesheet version="3.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"
  xmlns:dc="data:,dpc"
  exclude-result-prefixes="#all">
  
<xsl:output method="xml" omit-xml-declaration="yes" encoding="UTF-8" indent="yes" />

<xsl:import href="https://raw.githubusercontent.com/davidcarlisle/web-xslt/main/htmlparse/htmlparse.xsl"/>

<xsl:mode on-no-match="shallow-copy"/>

<xsl:template match="TEQuery">
    <TaxExpenditure>
      <xsl:attribute name="id" select="IntID"/>
      <xsl:attribute name="TAXSORT" select="TaxSort"/>
      <xsl:value-of select = "TEName"/>
      <xsl:apply-templates select="@* | node()" />
    </TaxExpenditure>
</xsl:template>

<xsl:template match="Description">
  <xsl:copy>
    <xsl:apply-templates select="dc:htmlparse(., '', true())"/>
  </xsl:copy>
</xsl:template>

<xsl:template match="li">
  <listitem>
    <xsl:value-of disable-output-escaping = "yes"   select="."/>
  </listitem>
</xsl:template>

<xsl:template match="idref">
  <idref>
   <xsl:apply-templates/>
  </idref>
</xsl:template>

<xsl:template match="ul[ul] | font | div">
  <xsl:apply-templates/>
</xsl:template>

<xsl:template match="ul[not(ul)]">
  <unorderedList>
    <xsl:apply-templates/>
  </unorderedList>
</xsl:template>

<xsl:template match="IntID"/>
<xsl:template match="TaxSort"/>
<xsl:template match="TEName"/>
</xsl:stylesheet>

方案分析与优化建议

现有方案的局限性

当前用disable-output-escaping(DOE)的方式虽能得到预期输出,但属于XSLT中的特殊手段,存在明显短板:

  • 兼容性受限:并非所有XSLT处理器都支持DOE,且Saxon在输出到DOM节点而非序列化文本时会忽略该设置;
  • 破坏节点树模型:将输出逻辑绑定到序列化环节,后续若要对<idref>节点做二次处理(如验证引用有效性、修改内容)会非常棘手;
  • 代码可读性差:维护者需要额外理解DOE的作用场景,增加了维护成本。

更优解决方案:修复双层转义问题

核心问题是Access导出时产生了双层转义:原始<idref>先被转义成&lt;idref&gt;,导出XML时又被再次转义成&amp;lt;idref&amp;gt;。正确的处理方式是在解析阶段就把双层转义还原成实际节点,而非在输出阶段强制取消转义。

修改后的关键代码

  1. 修复Description模板,先还原双层转义再解析:
<xsl:template match="Description">
  <xsl:copy>
    <!-- 一次性还原所有双层转义的HTML实体 -->
    <xsl:variable name="unescaped-text">
      <xsl:value-of select="." disable-output-escaping="yes"/>
    </xsl:variable>
    <!-- 解析为节点树 -->
    <xsl:apply-templates select="dc:htmlparse($unescaped-text, '', true())"/>
  </xsl:copy>
</xsl:template>
  1. 移除li模板中的DOE,改为正常应用模板:
<xsl:template match="li">
  <listitem>
    <xsl:apply-templates/>
  </listitem>
</xsl:template>

方案优势

  1. 符合XSLT核心模型:所有操作在节点树层面完成,后续可轻松对<idref>节点进行任何扩展操作;
  2. 兼容性更强:无需依赖DOE的处理器支持,在Saxon EE 9.8.3及其他主流处理器中都能稳定运行;
  3. 逻辑更清晰:直白处理双层转义问题,代码可读性和维护性大幅提升。

内容的提问来源于stack exchange,提问作者MadeFrame

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 16:50:54