如何将含嵌套标签的XML加载到DataTable并避免DuplicateNameException
解决DataSet加载XML时的DuplicateNameException问题及替代方案
一、修复DataSet加载的冲突问题
DataSet自动推断XML结构时,会将同名标签映射为同名列,导致重复列冲突。解决核心是手动定义XML Schema(XSD),明确指定每个数据表的列名和结构,给重复的<orgao>标签分配不同的列名。
步骤1:编写自定义XSD文件
假设你的XML核心结构如下:
<root> <listaNcm>...</listaNcm> <detalhesAtributos> <atributo> <orgao>主机构</orgao> <condicionados> <orgao>子机构</orgao> </condicionados> </atributo> </detalhesAtributos> </root>
对应的XSD需将<condicionados>内的<orgao>重命名为orgaoCondicionado,示例XSD:
<xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema"> <xs:element name="root"> <xs:complexType> <xs:sequence> <xs:element name="listaNcm" type="xs:string" minOccurs="0" maxOccurs="unbounded"/> <xs:element name="detalhesAtributos" maxOccurs="unbounded"> <xs:complexType> <xs:sequence> <xs:element name="atributo" maxOccurs="unbounded"> <xs:complexType> <xs:sequence> <xs:element name="orgao" type="xs:string"/> <xs:element name="condicionados"> <xs:complexType> <xs:sequence> <!-- 给子节点的orgao重命名 --> <xs:element name="orgao" type="xs:string" form="qualified" name="orgaoCondicionado"/> </xs:sequence> </xs:complexType> </xs:element> </xs:sequence> </xs:complexType> </xs:element> </xs:sequence> </xs:complexType> </xs:element> </xs:sequence> </xs:complexType> </xs:element> </xs:schema>
注意:需根据XML完整结构补全所有元素的定义。
步骤2:修改C#代码加载XSD和XML
先加载自定义XSD定义表结构,再读取XML,避免自动推断导致的列名冲突:
public string LerArq(string pathXML, string pathXSD) { DataSet dsXml = new DataSet(); string sResult = string.Empty; if (!string.IsNullOrEmpty(pathXML) && File.Exists(pathXML) && File.Exists(pathXSD)) { try { // 先加载自定义XSD确定表结构 dsXml.ReadXmlSchema(pathXSD); // 再加载XML数据 dsXml.ReadXml(pathXML); sResult = "XML加载成功"; // 后续可通过DataSet操作数据更新数据库 } catch (Exception e) { sResult = $"加载失败:{e.Message}"; } } return sResult; }
二、适合大XML文件的替代处理方式
你的XML有60万行,DataSet会将整个文件加载到内存,可能引发内存压力问题。以下是更高效的替代方案:
1. XmlReader(流式处理,低内存占用)
XmlReader是只读、向前的流式解析器,不会一次性加载整个XML到内存,适合超大文件。可逐节点读取并直接写入数据库:
public string ProcessLargeXml(string pathXML) { string sResult = string.Empty; if (!File.Exists(pathXML)) return "文件不存在"; using (XmlReader reader = XmlReader.Create(pathXML)) { try { while (reader.Read()) { if (reader.NodeType == XmlNodeType.Element) { switch (reader.Name) { case "atributo": // 读取主节点的orgao reader.ReadToDescendant("orgao"); string mainOrgao = reader.ReadElementContentAsString(); // 读取子节点condicionados下的orgao reader.ReadToFollowing("condicionados"); if (reader.ReadToDescendant("orgao")) { string condOrgao = reader.ReadElementContentAsString(); // 直接调用数据库操作写入数据 // InsertIntoDatabase(mainOrgao, condOrgao); } break; case "listaNcm": string ncmData = reader.ReadElementContentAsString(); // 处理listaNcm数据 break; } } } sResult = "XML处理完成"; } catch (Exception e) { sResult = $"处理失败:{e.Message}"; } } return sResult; }
2. LINQ to XML(XDocument)
若服务器内存充足(60万行XML约几十到上百MB),LINQ to XML提供更简洁的查询语法,方便提取数据:
public string ProcessXmlWithLinq(string pathXML) { string sResult = string.Empty; if (!File.Exists(pathXML)) return "文件不存在"; try { XDocument doc = XDocument.Load(pathXML); // 查询所有atributo节点,提取两个orgao的值 var atributos = from attr in doc.Descendants("atributo") select new { MainOrgao = attr.Element("orgao")?.Value, CondOrgao = attr.Descendants("condicionados")?.Elements("orgao")?.FirstOrDefault()?.Value }; // 遍历数据并更新数据库 foreach (var attr in atributos) { // InsertIntoDatabase(attr.MainOrgao, attr.CondOrgao); } sResult = "XML处理完成"; } catch (Exception e) { sResult = $"处理失败:{e.Message}"; } return sResult; }
注意:超大文件使用XDocument可能导致内存溢出,此时优先选择XmlReader。
三、总结
- 坚持用DataSet的话,必须通过自定义XSD明确表结构,规避列名重复;
- 大XML文件优先选XmlReader,内存占用低、性能稳定;
- 内存充足时,LINQ to XML操作更简洁,适合快速开发。
内容的提问来源于stack exchange,提问作者F_Bug
相关产品推荐
相关产品推荐

