You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Groovy移除XML Payload中的重复results节点?

问题:XML提取指定节点后移除重复的results组合

我尝试从XML Payload中提取指定节点(排除points、StartDate、EndDate),但生成的输出中出现了重复的results节点组合。请问是否有办法获取值的唯一组合,或是在后续步骤中移除重复值?

原Groovy代码

import java.text.*
import groovy.xml.*

def text = '''
<root>
  <results>
    <loc>Loc 10</loc>
    <city>ABC</city>
    <points>3</points>
    <StartDate>2023-09-11T22:39:40Z</StartDate>
    <EndDate>2023-09-13T22:45:36.437000Z</EndDate>
  </results>
  <results>
    <loc>Loc 11</loc>
   <city>ABC</city> 
    <points>4</points>
    <StartDate>2023-09-18T22:39:40Z</StartDate>
    <EndDate>2023-09-18T22:45:36.437000Z</EndDate>
  </results>
  <results>
    <loc>Loc 11</loc>
    <city>ABC</city>
    <points>4</points>
    <StartDate>2023-02-16T22:39:40Z</StartDate>
    <EndDate>2023-09-18T22:45:36.437000Z</EndDate>
  </results>
  <results>
    <loc>Loc 11</loc>
    <city>XYZ</city>
    <points>4</points>
    <StartDate>2023-09-16T22:39:40Z</StartDate>
    <EndDate>2023-12-18T22:45:36.437000Z</EndDate>
  </results>
</root>
'''
def xml = new XmlSlurper().parseText( text )
def output = new XmlParser().parseText("<root/>")

xml.results.each { resXml ->
       Node resultsNode = output.appendNode( new QName("results"), [:] )
       resXml.children().findAll { child -> child.name() != "points" && child.name()!= "StartDate" && child.name() != "EndDate" }.each { child ->
          resultsNode.appendNode( new QName(child.name()), [:], child.text() )
       }
}

println XmlUtil.serialize(output ) 

现有输出(含重复节点)

<?xml version="1.0" encoding="UTF-8"?><root>
    
  <results xmlns="">
        
    <loc>Loc 10</loc>
        
    <city>ABC</city>
      
  </results>
    
  <results xmlns="">
        
    <loc>Loc 11</loc>
        
    <city>ABC</city>
        
    <loc_name>Loc Desc 11</loc_name>
      
  </results>
    
  <results xmlns="">
        
    <loc>Loc 11</loc>
        
    <city>ABC</city>
      
  </results>
    
  <results xmlns="">
        
    <loc>Loc 11</loc>
        
    <city>XYZ</city>
      
  </results>
  
</root>

解决方案

可以通过两种方式实现去重:将节点内容转为Map后批量去重,或者在遍历过程中记录已添加的组合避免重复。

方法一:利用Map去重

将每个results节点的有效内容转为键值对Map,利用Groovy的unique()方法自动去重(相同键值的Map会被判定为相等),再重新构建XML节点:

import java.text.*
import groovy.xml.*

def text = '''
<root>
  <results>
    <loc>Loc 10</loc>
    <city>ABC</city>
    <points>3</points>
    <StartDate>2023-09-11T22:39:40Z</StartDate>
    <EndDate>2023-09-13T22:45:36.437000Z</EndDate>
  </results>
  <results>
    <loc>Loc 11</loc>
   <city>ABC</city> 
    <points>4</points>
    <StartDate>2023-09-18T22:39:40Z</StartDate>
    <EndDate>2023-09-18T22:45:36.437000Z</EndDate>
  </results>
  <results>
    <loc>Loc 11</loc>
    <city>ABC</city>
    <points>4</points>
    <StartDate>2023-02-16T22:39:40Z</StartDate>
    <EndDate>2023-09-18T22:45:36.437000Z</EndDate>
  </results>
  <results>
    <loc>Loc 11</loc>
    <city>XYZ</city>
    <points>4</points>
    <StartDate>2023-09-16T22:39:40Z</StartDate>
    <EndDate>2023-12-18T22:45:36.437000Z</EndDate>
  </results>
</root>
'''
def xml = new XmlSlurper().parseText(text)
def output = new XmlParser().parseText("<root/>")

// 提取有效节点并转为Map,再去重
def uniqueResults = xml.results.collect { resXml ->
    def filteredNodes = resXml.children().findAll { child ->
        !['points', 'StartDate', 'EndDate'].contains(child.name())
    }
    filteredNodes.collectEntries { [(it.name()): it.text()] }
}.unique()

// 将去重后的Map转换为XML节点
uniqueResults.each { resultMap ->
    Node resultsNode = output.appendNode(new QName("results"), [:])
    resultMap.each { key, value ->
        resultsNode.appendNode(new QName(key), [:], value)
    }
}

println XmlUtil.serialize(output)

方法二:遍历过程中记录已添加组合

用一个Set存储已处理过的节点组合的唯一标识,每次处理前检查标识是否存在,不存在才添加到输出XML:

import java.text.*
import groovy.xml.*

def text = '''
<root>
  <results>
    <loc>Loc 10</loc>
    <city>ABC</city>
    <points>3</points>
    <StartDate>2023-09-11T22:39:40Z</StartDate>
    <EndDate>2023-09-13T22:45:36.437000Z</EndDate>
  </results>
  <results>
    <loc>Loc 11</loc>
   <city>ABC</city> 
    <points>4</points>
    <StartDate>2023-09-18T22:39:40Z</StartDate>
    <EndDate>2023-09-18T22:45:36.437000Z</EndDate>
  </results>
  <results>
    <loc>Loc 11</loc>
    <city>ABC</city>
    <points>4</points>
    <StartDate>2023-02-16T22:39:40Z</StartDate>
    <EndDate>2023-09-18T22:45:36.437000Z</EndDate>
  </results>
  <results>
    <loc>Loc 11</loc>
    <city>XYZ</city>
    <points>4</points>
    <StartDate>2023-09-16T22:39:40Z</StartDate>
    <EndDate>2023-12-18T22:45:36.437000Z</EndDate>
  </results>
</root>
'''
def xml = new XmlSlurper().parseText(text)
def output = new XmlParser().parseText("<root/>")
def seenCombinations = new HashSet<String>()

xml.results.each { resXml ->
    def filteredNodes = resXml.children().findAll { child ->
        !['points', 'StartDate', 'EndDate'].contains(child.name())
    }
    // 生成唯一标识:按节点名排序后拼接键值,避免顺序不同导致误判
    def comboKey = filteredNodes.sort { it.name() }.collect { "${it.name()}:${it.text()}" }.join("|")
    if (!seenCombinations.contains(comboKey)) {
        seenCombinations.add(comboKey)
        Node resultsNode = output.appendNode(new QName("results"), [:])
        filteredNodes.each { child ->
            resultsNode.appendNode(new QName(child.name()), [:], child.text())
        }
    }
}

println XmlUtil.serialize(output)

两种方法最终都会输出不含重复results节点的XML。


内容的提问来源于stack exchange,提问作者N21RL

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 23:02:33