如何使用XPath将HTML拆分转换为DocBook XML的simplesect章节?
解决HTML转DocBook时按
拆分的XPath问题
要实现将HTML中每个<h3>及其后续关联内容(段落、figure等)转换为DocBook的<simplesect>,核心是精准选中每个<h3>对应的后续兄弟节点,直到下一个<h3>(或父容器末尾)。以下分XPath版本给出可行表达式:
XPath 1.0 适配方案
由于XPath 1.0不支持原生区间选择,需通过节点计数界定范围。假设当前上下文为目标<h3>节点,选中其对应内容的表达式为:
following-sibling::*[not(self::h3) and count(preceding-sibling::h3) = count(current()/preceding-sibling::h3)]
表达式解释:
following-sibling::*:选中当前<h3>之后的所有兄弟元素not(self::h3):排除后续的<h3>节点,避免混入下一个章节标题count(preceding-sibling::h3) = count(current()/preceding-sibling::h3):确保目标节点的前置<h3>数量与当前<h3>的前置数量一致,以此锁定属于当前章节的内容范围
若要同时选中<h3>本身和对应内容,可合并表达式:
self::h3 | following-sibling::*[not(self::h3) and count(preceding-sibling::h3) = count(current()/preceding-sibling::h3)]
XPath 2.0+ 简化方案
XPath 2.0及以上支持更直观的节点区间语法,可直接定位到下一个<h3>之前的内容:
following-sibling::*[not(self::h3) and (. >> current() and (not(following-sibling::h3) or . << following-sibling::h3[1]))]
表达式解释:
. >> current():确保节点在当前<h3>之后not(following-sibling::h3) or . << following-sibling::h3[1]:要么没有后续<h3>(已是最后一个章节),要么节点在第一个后续<h3>之前
实际XSLT转换示例
结合XSLT实现完整转换(以XPath 1.0为例),假设HTML结构如下:
<div class="product-content"> <h3>核心功能</h3> <p>产品支持多终端同步,满足跨场景使用需求。</p> <figure><img src="sync.png" alt="同步示意图"></figure> <h3>快速上手</h3> <p>注册账号后,完成三步配置即可启用服务:</p> <ol> <li>绑定设备</li> <li>设置同步规则</li> <li>启动同步任务</li> </ol> </div>
对应的XSLT模板:
<xsl:template match="div[@class='product-content']"> <chapter> <xsl:for-each select="h3"> <simplesect> <title><xsl:value-of select="text()"/></title> <!-- 复制当前<h3>对应的所有内容节点 --> <xsl:copy-of select="following-sibling::*[not(self::h3) and count(preceding-sibling::h3) = count(current()/preceding-sibling::h3)]"/> </simplesect> </xsl:for-each> </chapter> </xsl:template>
内容的提问来源于stack exchange,提问作者web_ou
相关产品推荐
相关产品推荐

