You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Groovy根据特殊字符第N次出现拆分多XML标签?

问题描述

我有一个包含大量数据的多标签XML输入,希望使用Groovy根据输入标签中特殊字符的第N次出现来拆分XML标签。参考过Java中「按字符第N次出现拆分字符串」的逻辑,但该逻辑仅适用于单个XML标签,我需要对所有标签(例如按分隔符/的第4次出现)进行统一拆分处理。

示例输入XML

<Row>
<EntityID>9035158701/9035158702/9035158703/9035158704/9035158705/9035158706/9035158707/9035158708/9035158709/9035158710</EntityID>
<RefID>7U2HYUTP/5Z1IWGUS/7AK9MDDJ/6RP9DXAW/29FBRBEL/5YKDCO3B/75MUQU7S/57QCGOQE/2EUX64ON/2VTJPVUV</RefID>
</Row>

期望输出XML

<Root>
<Row>
<EntityID>9035158701/9035158702/9035158703/9035158704</EntityID>
<RefID>7U2HYUTP/5Z1IWGUS/7AK9MDDJ/6RP9DXAW</RefID>
</Row>
<Row>
<EntityID>9035158705/9035158706/9035158707/9035158708</EntityID>
<RefID>29FBRBEL/5YKDCO3B/75MUQU7S/57QCGOQE</RefID>
</Row>
<Row>
<EntityID>9035158709/9035158710</EntityID>
<RefID>2EUX64ON/2VTJPVUV</RefID>
</Row>
</Root>

现有代码(未完成多标签处理)

import com.sap.gateway.ip.core.customdev.util.Message;
import java.util.HashMap;
import groovy.xml.*;

def Message processData(Message message) {

        def body = message.getBody(java.io.Reader)
        assert body != null
        def input = new XmlSlurper().parse(body);
        def entityid = input.EntityID.text();
        def refid = input.RefID.text();
        int nth=0;
        int cont=0;
        
        def writer = new StringWriter()
        for(int i=0;i<entityid.length();i++){
            if(entityid.charAt(i)=='/')
                nth++;
                
            if(nth == 4 || i==entityid.length()-1){
                new MarkupBuilder(writer).Row{
                
                if(i==entityid.length()-1) //with this if you preveent to cut the last number
                EntityID (entityid.substring(cont,i+1))
                
                else
                    EntityID (entityid.substring(cont,i+1))
                nth=0;
                cont =i+1;
            }
        }
        }
       def output = writer.toString()
       message.setBody(output)      
       return message;
}
实现方案

核心思路

将每个标签的内容按分隔符拆分为元素列表,按每组4个元素(对应分隔符/的第4次出现)进行分组,再将各组内容拼接后生成对应的XML结构,确保所有标签的分组位置完全对齐。

完整代码

import com.sap.gateway.ip.core.customdev.util.Message;
import groovy.xml.MarkupBuilder;

def Message processData(Message message) {
    def body = message.getBody(java.io.Reader)
    assert body != null
    def input = new XmlSlurper().parse(body)
    
    // 定义每组包含的元素数量(4个元素对应3个/,即第4次出现分隔符时拆分)
    final int GROUP_SIZE = 4
    
    // 拆分各标签内容为列表
    def entityList = input.EntityID.text().split('/') as List
    def refList = input.RefID.text().split('/') as List
    
    // 按指定大小分组,最后一组可能不足GROUP_SIZE
    def entityGroups = entityList.collate(GROUP_SIZE)
    def refGroups = refList.collate(GROUP_SIZE)
    
    // 生成输出XML
    def writer = new StringWriter()
    def builder = new MarkupBuilder(writer)
    builder.Root {
        // 遍历分组,同步生成每个Row
        entityGroups.eachWithIndex { group, index ->
            Row {
                EntityID(group.join('/'))
                RefID(refGroups[index].join('/'))
            }
        }
    }
    
    message.setBody(writer.toString())
    return message
}

关键说明

  1. 拆分与分组:使用split('/')将标签内容拆分为单个元素的列表,再用Groovy内置的collate(GROUP_SIZE)方法按指定大小分组,自动处理最后一组不足数量的情况。
  2. 同步分组:确保EntityID和RefID的分组索引一一对应,保证每个<Row>内的内容是匹配的。
  3. XML生成:用MarkupBuilder规范生成XML结构,自动处理标签闭合,避免手动拼接字符串的错误。

内容的提问来源于stack exchange,提问作者Kavitha Sivaprakasam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 04:55:15