You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

处理大型XML文件重复UCI元素,合并关联Contact节点

解决方案

Python 实现(适配大文件,内存友好)

用xml.etree.ElementTree处理,核心逻辑是按UCI分组存储节点,合并后重新生成XML:

import xml.etree.ElementTree as ET
from collections import defaultdict

# 解析输入XML
tree = ET.parse('input.xml')
root = tree.getroot()

# 按UCI分组,存储基础client节点和待合并的Contact列表
uci_groups = defaultdict(lambda: {'client': None, 'contacts': []})

# 遍历所有client节点
for client in root.findall('client'):
    uci_node = client.find('UCI')
    if not uci_node:
        continue
    uci = uci_node.text
    contact = client.find('Contact')
    
    if not uci_groups[uci]['client']:
        # 保留第一个client节点作为合并基础
        uci_groups[uci]['client'] = client
    # 收集Contact节点
    if contact:
        uci_groups[uci]['contacts'].append(contact)

# 清空原根节点下的所有内容
for child in list(root):
    root.remove(child)

# 将合并后的节点添加回根节点
for group in uci_groups.values():
    client = group['client']
    for contact in group['contacts']:
        client.append(contact)
    root.append(client)

# 保存处理后的XML
tree.write('output.xml', encoding='utf-8', xml_declaration=True)

注:如果XML包含命名空间,需要在find/findall时带上命名空间前缀;若文件极大,建议改用ET.iterparse迭代解析,减少内存占用。

C# 实现(高效处理超大文件)

采用XmlDocument实现内存内合并,若文件超出内存承载能力,可改用XmlReader+XmlWriter流式处理:

using System;
using System.Collections.Generic;
using System.Xml;

class XmlUciMerger
{
    static void Main(string[] args)
    {
        string inputPath = "input.xml";
        string outputPath = "output.xml";
        Dictionary<string, bool> uciRecorded = new Dictionary<string, bool>();
        Dictionary<string, List<XmlNode>> uciContacts = new Dictionary<string, List<XmlNode>>();

        XmlDocument doc = new XmlDocument();
        doc.Load(inputPath);
        XmlNode root = doc.DocumentElement;

        // 遍历所有client节点,分组处理
        foreach (XmlNode client in root.SelectNodes("client"))
        {
            XmlNode uciNode = client.SelectSingleNode("UCI");
            if (uciNode == null) continue;
            string uci = uciNode.InnerText;
            XmlNode contactNode = client.SelectSingleNode("Contact");

            if (!uciRecorded.ContainsKey(uci))
            {
                uciRecorded.Add(uci, true);
                if (contactNode != null)
                {
                    if (!uciContacts.ContainsKey(uci))
                        uciContacts.Add(uci, new List<XmlNode>());
                    uciContacts[uci].Add(contactNode.Clone());
                }
            }
            else
            {
                // 收集当前节点的Contact,然后删除重复的client节点
                if (contactNode != null)
                {
                    if (!uciContacts.ContainsKey(uci))
                        uciContacts.Add(uci, new List<XmlNode>());
                    uciContacts[uci].Add(contactNode.Clone());
                }
                root.RemoveChild(client);
            }
        }

        // 将收集的Contact追加到对应client节点下
        foreach (var kvp in uciContacts)
        {
            XmlNode targetClient = root.SelectSingleNode($"client[UCI='{kvp.Key}']");
            if (targetClient != null)
            {
                foreach (XmlNode contact in kvp.Value)
                {
                    targetClient.AppendChild(contact);
                }
            }
        }

        doc.Save(outputPath);
    }
}

注:流式处理版本可通过XmlReader逐节点读取,遇到首个UCI的client时用XmlWriter写入节点开头,后续同一UCI的节点仅提取Contact暂存,最后统一写入对应client节点后。

内容的提问来源于stack exchange,提问作者Baron89

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 19:35:25