XSLT技术实现:提取Word索引条目为连续字符串并保留非索引相邻元素
解决Word XML索引条目提取并保留原有内容的问题
我明白你之前的困扰:想用for-each-group处理索引条目,结果不小心把非索引的文本都弄丢了。咱们换个更稳妥的思路——用身份模板保留所有默认内容,再专门针对索引节点做替换处理,这样既能提取索引内容,又不会破坏原有文档结构。
核心解决方案思路
- 身份模板(Identity Template):作为基础逻辑,自动复制所有不需要修改的节点和属性,确保普通文本内容完整保留。
- 精准定位索引边界:找到索引起始节点
w:r[w:fldChar[@w:fldCharType='begin']],并定位到对应的结束节点w:r[w:fldChar[@w:fldCharType='end']]。 - 合并索引内容:提取起止节点之间所有
w:instrText的文本,合并成连续字符串,替换为<indexentry>元素。 - 跳过原索引节点:避免原有的索引起止节点和中间的
w:r被重复输出。
完整XSLT代码
<xsl:stylesheet version="2.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform" xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main"> <!-- 身份模板:复制所有未被自定义匹配的节点/属性 --> <xsl:template match="@*|node()"> <xsl:copy> <xsl:apply-templates select="@*|node()"/> </xsl:copy> </xsl:template> <!-- 匹配索引开始的w:r节点 --> <xsl:template match="w:r[w:fldChar[@w:fldCharType='begin']]"> <!-- 定位对应的索引结束节点 --> <xsl:variable name="end-node" select="following-sibling::w:r[w:fldChar[@w:fldCharType='end']][1]"/> <!-- 提取并合并中间所有instrText的文本内容 --> <indexentry> <xsl:value-of select="following-sibling::w:r[. << $end-node]/w:instrText/text()" separator=""/> </indexentry> <!-- 跳过结束节点的输出 --> <xsl:apply-templates select="$end-node" mode="skip"/> </xsl:template> <!-- 跳过索引结束节点的输出 --> <xsl:template match="w:r[w:fldChar[@w:fldCharType='end']]" mode="skip"/> <!-- 跳过索引起止节点之间的w:r节点(避免被身份模板重复复制) --> <xsl:template match="w:r[preceding-sibling::w:r[w:fldChar[@w:fldCharType='begin']] and following-sibling::w:r[w:fldChar[@w:fldCharType='end']]]"/> </xsl:stylesheet>
代码细节解释
- 身份模板:这是XSLT中处理“保留大部分内容”的标准方案,确保普通的
w:r、w:t等节点都会被原样复制,不会丢失。 - 索引内容提取:用
following-sibling::w:r[. << $end-node]精准选中起止节点之间的所有w:r,再通过value-of合并所有w:instrText的文本,得到连续的索引字符串。 - 节点跳过逻辑:通过专门的模板跳过索引起止节点和中间的
w:r,避免这些原节点被重复输出。
测试输出结果
用你提供的源XML运行上述代码后,会得到完全符合预期的输出:
<w:p xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main"> <w:r> <w:t>Stahl und Beton</w:t> </w:r> <w:r> <w:t>lässt</w:t> </w:r> <w:r> <w:t> sich aus dem alten Namen des Materials ableiten.</w:t> </w:r> <indexentry> XE "Moniereisen" </indexentry> </w:p>
为什么之前的代码出问题?
你之前的模板只聚焦于处理包含索引的w:p,并且在for-each-group里只输出了索引相关内容,完全没有处理和复制普通的w:r节点,自然会导致非索引文本丢失。而身份模板+精准匹配索引节点的方式,完美解决了这个矛盾。
内容的提问来源于stack exchange,提问作者Pjoern
相关产品推荐
相关产品推荐

