如何用Groovy移除XML Payload中的重复results节点?
问题:XML提取指定节点后移除重复的results组合
我尝试从XML Payload中提取指定节点(排除points、StartDate、EndDate),但生成的输出中出现了重复的results节点组合。请问是否有办法获取值的唯一组合,或是在后续步骤中移除重复值?
原Groovy代码
import java.text.* import groovy.xml.* def text = ''' <root> <results> <loc>Loc 10</loc> <city>ABC</city> <points>3</points> <StartDate>2023-09-11T22:39:40Z</StartDate> <EndDate>2023-09-13T22:45:36.437000Z</EndDate> </results> <results> <loc>Loc 11</loc> <city>ABC</city> <points>4</points> <StartDate>2023-09-18T22:39:40Z</StartDate> <EndDate>2023-09-18T22:45:36.437000Z</EndDate> </results> <results> <loc>Loc 11</loc> <city>ABC</city> <points>4</points> <StartDate>2023-02-16T22:39:40Z</StartDate> <EndDate>2023-09-18T22:45:36.437000Z</EndDate> </results> <results> <loc>Loc 11</loc> <city>XYZ</city> <points>4</points> <StartDate>2023-09-16T22:39:40Z</StartDate> <EndDate>2023-12-18T22:45:36.437000Z</EndDate> </results> </root> ''' def xml = new XmlSlurper().parseText( text ) def output = new XmlParser().parseText("<root/>") xml.results.each { resXml -> Node resultsNode = output.appendNode( new QName("results"), [:] ) resXml.children().findAll { child -> child.name() != "points" && child.name()!= "StartDate" && child.name() != "EndDate" }.each { child -> resultsNode.appendNode( new QName(child.name()), [:], child.text() ) } } println XmlUtil.serialize(output )
现有输出(含重复节点)
<?xml version="1.0" encoding="UTF-8"?><root> <results xmlns=""> <loc>Loc 10</loc> <city>ABC</city> </results> <results xmlns=""> <loc>Loc 11</loc> <city>ABC</city> <loc_name>Loc Desc 11</loc_name> </results> <results xmlns=""> <loc>Loc 11</loc> <city>ABC</city> </results> <results xmlns=""> <loc>Loc 11</loc> <city>XYZ</city> </results> </root>
解决方案
可以通过两种方式实现去重:将节点内容转为Map后批量去重,或者在遍历过程中记录已添加的组合避免重复。
方法一:利用Map去重
将每个results节点的有效内容转为键值对Map,利用Groovy的unique()方法自动去重(相同键值的Map会被判定为相等),再重新构建XML节点:
import java.text.* import groovy.xml.* def text = ''' <root> <results> <loc>Loc 10</loc> <city>ABC</city> <points>3</points> <StartDate>2023-09-11T22:39:40Z</StartDate> <EndDate>2023-09-13T22:45:36.437000Z</EndDate> </results> <results> <loc>Loc 11</loc> <city>ABC</city> <points>4</points> <StartDate>2023-09-18T22:39:40Z</StartDate> <EndDate>2023-09-18T22:45:36.437000Z</EndDate> </results> <results> <loc>Loc 11</loc> <city>ABC</city> <points>4</points> <StartDate>2023-02-16T22:39:40Z</StartDate> <EndDate>2023-09-18T22:45:36.437000Z</EndDate> </results> <results> <loc>Loc 11</loc> <city>XYZ</city> <points>4</points> <StartDate>2023-09-16T22:39:40Z</StartDate> <EndDate>2023-12-18T22:45:36.437000Z</EndDate> </results> </root> ''' def xml = new XmlSlurper().parseText(text) def output = new XmlParser().parseText("<root/>") // 提取有效节点并转为Map,再去重 def uniqueResults = xml.results.collect { resXml -> def filteredNodes = resXml.children().findAll { child -> !['points', 'StartDate', 'EndDate'].contains(child.name()) } filteredNodes.collectEntries { [(it.name()): it.text()] } }.unique() // 将去重后的Map转换为XML节点 uniqueResults.each { resultMap -> Node resultsNode = output.appendNode(new QName("results"), [:]) resultMap.each { key, value -> resultsNode.appendNode(new QName(key), [:], value) } } println XmlUtil.serialize(output)
方法二:遍历过程中记录已添加组合
用一个Set存储已处理过的节点组合的唯一标识,每次处理前检查标识是否存在,不存在才添加到输出XML:
import java.text.* import groovy.xml.* def text = ''' <root> <results> <loc>Loc 10</loc> <city>ABC</city> <points>3</points> <StartDate>2023-09-11T22:39:40Z</StartDate> <EndDate>2023-09-13T22:45:36.437000Z</EndDate> </results> <results> <loc>Loc 11</loc> <city>ABC</city> <points>4</points> <StartDate>2023-09-18T22:39:40Z</StartDate> <EndDate>2023-09-18T22:45:36.437000Z</EndDate> </results> <results> <loc>Loc 11</loc> <city>ABC</city> <points>4</points> <StartDate>2023-02-16T22:39:40Z</StartDate> <EndDate>2023-09-18T22:45:36.437000Z</EndDate> </results> <results> <loc>Loc 11</loc> <city>XYZ</city> <points>4</points> <StartDate>2023-09-16T22:39:40Z</StartDate> <EndDate>2023-12-18T22:45:36.437000Z</EndDate> </results> </root> ''' def xml = new XmlSlurper().parseText(text) def output = new XmlParser().parseText("<root/>") def seenCombinations = new HashSet<String>() xml.results.each { resXml -> def filteredNodes = resXml.children().findAll { child -> !['points', 'StartDate', 'EndDate'].contains(child.name()) } // 生成唯一标识:按节点名排序后拼接键值,避免顺序不同导致误判 def comboKey = filteredNodes.sort { it.name() }.collect { "${it.name()}:${it.text()}" }.join("|") if (!seenCombinations.contains(comboKey)) { seenCombinations.add(comboKey) Node resultsNode = output.appendNode(new QName("results"), [:]) filteredNodes.each { child -> resultsNode.appendNode(new QName(child.name()), [:], child.text()) } } } println XmlUtil.serialize(output)
两种方法最终都会输出不含重复results节点的XML。
内容的提问来源于stack exchange,提问作者N21RL
相关产品推荐
相关产品推荐

