如何合并XML文档中连续的<bold>元素?
合并XML中连续
<bold>元素的解决方案 我们需要处理XML里的连续<bold>元素,只有当两个<bold>之间没有其他节点(不管有没有空格)时才合并它们;如果中间有文本或者其他元素,就保持各自独立。以下是具体的示例和实现代码:
输入XML
<?xml version="1.0" encoding="UTF-8"?> <root> <h1><bold>Abandonment of Trade Secret</bold></h1> <h2><bold>Abuse of Discretion.</bold> See <bold>Discretion of Court</bold></h2> <h3>Licensing agreement, modifications to. See <bold>Licensing Agreements</bold></h3> <h4><bold>Alternative Minimum Tax (AM</bold><bold>T)</bold></h4> <h5><bold>Audits.</bold> See <bold>Trade Secre</bold><bold>t Audits</bold></h5> <h6><bold>California Uniform Trade</bold><bold> Secrets Act (UTSA).</bold> See <bold>Uniform Trade Secrets Act, California (UTSA)</bold></h6> <h7><bold>Charts, Checklists, Questionnaires, </bold><bold>and Tables</bold></h7> <h8><bold>Competition.</bold> See <bold>Covenant Against Competition;</bold> <bold>Unfair Competition</bold></h8> <h9>See also <bold>Internet;</bold> <bold>Websites</bold> <bold>URL</bold></h9> <h10>See also <bold>Copyrights;</bold><bold> Intellectual Property; Patents; </bold><bold>Trademarks</bold></h10> </root>
预期输出XML
<?xml version="1.0" encoding="UTF-8"?> <root> <h1><bold>Abandonment of Trade Secret</bold></h1> <h2><bold>Abuse of Discretion.</bold> See <bold>Discretion of Court</bold></h2> <h3>Licensing agreement, modifications to. See <bold>Licensing Agreements</bold></h3> <h4><bold>Alternative Minimum Tax (AMT)</bold></h4> <h5><bold>Audits.</bold> See <bold>Trade Secret Audits</bold></h5> <h6><bold>California Uniform Trade Secrets Act (UTSA).</bold> See <bold>Uniform Trade Secrets Act, California (UTSA)</bold></h6> <h7><bold>Charts, Checklists, Questionnaires, and Tables</bold></h7> <h8><bold>Competition.</bold> See <bold>Covenant Against Competition; Unfair Competition</bold></h8> <h9>See also <bold>Internet; Websites URL</bold></h9> <h10>See also <bold>Copyrights; Intellectual Property; Patents; Trademarks</bold></h10> </root>
XSLT实现代码
<?xml version="1.0" encoding="UTF-8"?> <xsl:stylesheet xmlns:xsl="http://www.w3.org/1999/XSL/Transform" xmlns:xs="http://www.w3.org/2001/XMLSchema" exclude-result-prefixes="xs" version="2.0"> <xsl:template match="node()|@*"> <xsl:copy> <xsl:apply-templates select="node()|@*"/> </xsl:copy> </xsl:template> <xsl:template match="*[bold]"> <xsl:copy> <xsl:for-each-group select="*" group-adjacent=". instance of element(bold)"> <xsl:choose> <xsl:when test="current-grouping-key()"> <xsl:copy> <xsl:apply-templates select="current-group()/node()"/> </xsl:copy> </xsl:when> <xsl:otherwise> <xsl:apply-templates select="current-group()"/> </xsl:otherwise> </xsl:choose> </xsl:for-each-group> </xsl:copy> </xsl:template> </xsl:stylesheet>
代码说明
- 身份模板:第一个模板
match="node()|@*"会原样复制所有节点和属性,确保XML中不需要处理的部分保持不变。 - 分组处理模板:第二个模板针对所有包含
<bold>元素的父节点,通过group-adjacent=". instance of element(bold)"将相邻的<bold>元素归为一组,非<bold>元素单独成组。 - 合并逻辑:对于
<bold>元素组,我们创建一个新的<bold>元素,把组内所有<bold>的子节点都嵌入进去,实现内容合并;非<bold>组则直接输出,不做修改。
内容的提问来源于stack exchange,提问作者Umaima Fatima
相关产品推荐
相关产品推荐

