You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将含嵌套标签的XML加载到DataTable并避免DuplicateNameException

解决DataSet加载XML时的DuplicateNameException问题及替代方案

一、修复DataSet加载的冲突问题

DataSet自动推断XML结构时,会将同名标签映射为同名列,导致重复列冲突。解决核心是手动定义XML Schema(XSD),明确指定每个数据表的列名和结构,给重复的<orgao>标签分配不同的列名。

步骤1:编写自定义XSD文件

假设你的XML核心结构如下:

<root>
  <listaNcm>...</listaNcm>
  <detalhesAtributos>
    <atributo>
      <orgao>主机构</orgao>
      <condicionados>
        <orgao>子机构</orgao>
      </condicionados>
    </atributo>
  </detalhesAtributos>
</root>

对应的XSD需将<condicionados>内的<orgao>重命名为orgaoCondicionado,示例XSD:

<xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema">
  <xs:element name="root">
    <xs:complexType>
      <xs:sequence>
        <xs:element name="listaNcm" type="xs:string" minOccurs="0" maxOccurs="unbounded"/>
        <xs:element name="detalhesAtributos" maxOccurs="unbounded">
          <xs:complexType>
            <xs:sequence>
              <xs:element name="atributo" maxOccurs="unbounded">
                <xs:complexType>
                  <xs:sequence>
                    <xs:element name="orgao" type="xs:string"/>
                    <xs:element name="condicionados">
                      <xs:complexType>
                        <xs:sequence>
                          <!-- 给子节点的orgao重命名 -->
                          <xs:element name="orgao" type="xs:string" form="qualified" name="orgaoCondicionado"/>
                        </xs:sequence>
                      </xs:complexType>
                    </xs:element>
                  </xs:sequence>
                </xs:complexType>
              </xs:element>
            </xs:sequence>
          </xs:complexType>
        </xs:element>
      </xs:sequence>
    </xs:complexType>
  </xs:element>
</xs:schema>

注意:需根据XML完整结构补全所有元素的定义。

步骤2:修改C#代码加载XSD和XML

先加载自定义XSD定义表结构,再读取XML,避免自动推断导致的列名冲突:

public string LerArq(string pathXML, string pathXSD)
{
    DataSet dsXml = new DataSet();
    string sResult = string.Empty;

    if (!string.IsNullOrEmpty(pathXML) && File.Exists(pathXML) && File.Exists(pathXSD))
    {
        try
        {
            // 先加载自定义XSD确定表结构
            dsXml.ReadXmlSchema(pathXSD);
            // 再加载XML数据
            dsXml.ReadXml(pathXML);
            sResult = "XML加载成功";
            // 后续可通过DataSet操作数据更新数据库
        }
        catch (Exception e)
        {
            sResult = $"加载失败:{e.Message}";
        }
    }

    return sResult;
}

二、适合大XML文件的替代处理方式

你的XML有60万行,DataSet会将整个文件加载到内存,可能引发内存压力问题。以下是更高效的替代方案:

1. XmlReader(流式处理,低内存占用)

XmlReader是只读、向前的流式解析器,不会一次性加载整个XML到内存,适合超大文件。可逐节点读取并直接写入数据库:

public string ProcessLargeXml(string pathXML)
{
    string sResult = string.Empty;
    if (!File.Exists(pathXML)) return "文件不存在";

    using (XmlReader reader = XmlReader.Create(pathXML))
    {
        try
        {
            while (reader.Read())
            {
                if (reader.NodeType == XmlNodeType.Element)
                {
                    switch (reader.Name)
                    {
                        case "atributo":
                            // 读取主节点的orgao
                            reader.ReadToDescendant("orgao");
                            string mainOrgao = reader.ReadElementContentAsString();
                            
                            // 读取子节点condicionados下的orgao
                            reader.ReadToFollowing("condicionados");
                            if (reader.ReadToDescendant("orgao"))
                            {
                                string condOrgao = reader.ReadElementContentAsString();
                                // 直接调用数据库操作写入数据
                                // InsertIntoDatabase(mainOrgao, condOrgao);
                            }
                            break;
                        case "listaNcm":
                            string ncmData = reader.ReadElementContentAsString();
                            // 处理listaNcm数据
                            break;
                    }
                }
            }
            sResult = "XML处理完成";
        }
        catch (Exception e)
        {
            sResult = $"处理失败:{e.Message}";
        }
    }
    return sResult;
}

2. LINQ to XML(XDocument)

若服务器内存充足(60万行XML约几十到上百MB),LINQ to XML提供更简洁的查询语法,方便提取数据:

public string ProcessXmlWithLinq(string pathXML)
{
    string sResult = string.Empty;
    if (!File.Exists(pathXML)) return "文件不存在";

    try
    {
        XDocument doc = XDocument.Load(pathXML);
        
        // 查询所有atributo节点,提取两个orgao的值
        var atributos = from attr in doc.Descendants("atributo")
                        select new
                        {
                            MainOrgao = attr.Element("orgao")?.Value,
                            CondOrgao = attr.Descendants("condicionados")?.Elements("orgao")?.FirstOrDefault()?.Value
                        };
        
        // 遍历数据并更新数据库
        foreach (var attr in atributos)
        {
            // InsertIntoDatabase(attr.MainOrgao, attr.CondOrgao);
        }
        
        sResult = "XML处理完成";
    }
    catch (Exception e)
    {
        sResult = $"处理失败:{e.Message}";
    }
    return sResult;
}

注意:超大文件使用XDocument可能导致内存溢出,此时优先选择XmlReader。

三、总结

  • 坚持用DataSet的话,必须通过自定义XSD明确表结构,规避列名重复;
  • 大XML文件优先选XmlReader,内存占用低、性能稳定;
  • 内存充足时,LINQ to XML操作更简洁,适合快速开发。

内容的提问来源于stack exchange,提问作者F_Bug

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.15 13:10:57