You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在C#中解析具有可变子节点结构的XML文件

嘿,我来帮你搞定这个XML解析的事儿!先把你给出的XML内容整理成格式化的代码块,然后给你几个实用的解析方案,不管你用Python还是其他工具,都能轻松上手:

目标XML结构(格式化后)
<?xml version="1.0" encoding="windows-1252" ?>
<TABLE>
  <COMPO>
    <alim_code> 1000 </alim_code>
    <const_code> 56700 </const_code>
    <teneur> 0 </teneur>
    <min missing=" " />
    <max missing=" " />
    <code_confiance> D </code_confiance>
    <source_code missing="." />
  </COMPO>
  <COMPO>
    <alim_code> 1000 </alim_code>
    <const_code> 60000 &amp;... <!-- 注意:XML中&必须转义为&amp;,否则会触发解析错误 -->
  </COMPO>
</TABLE>
解析实用方案

1. 先解决XML合法性问题

你提供的XML里有未转义的&,这是XML语法的硬伤——特殊字符(&、<、>等)必须转义。如果没法修改原始文件,建议用容错性强的解析库(比如Python的lxml)来绕过这个问题。

2. Python内置库解析(xml.etree.ElementTree)

这是无需额外安装的原生方案,适合快速开发:

import xml.etree.ElementTree as ET
import codecs

# 读取windows-1252编码的XML文件
with codecs.open('your_xml_file.xml', 'r', 'windows-1252') as f:
    xml_content = f.read()
    # 手动修复未转义的&(根据实际情况调整,避免误改合法内容)
    xml_content = xml_content.replace('&', '&amp;')

# 解析XML根节点
root = ET.fromstring(xml_content)

# 遍历所有COMPO节点
for compo in root.findall('COMPO'):
    # 提取普通文本节点(处理空值和多余空格)
    alim_code = compo.find('alim_code').text.strip() if compo.find('alim_code').text else None
    const_code = compo.find('const_code').text.strip() if compo.find('const_code').text else None
    teneur = compo.find('teneur').text.strip() if compo.find('teneur').text else None
    code_confiance = compo.find('code_confiance').text.strip() if compo.find('code_confiance').text else None

    # 提取带missing属性的节点
    min_missing = compo.find('min').get('missing') if compo.find('min') is not None else None
    max_missing = compo.find('max').get('missing') if compo.find('max') is not None else None
    source_missing = compo.find('source_code').get('missing') if compo.find('source_code') is not None else None

    # 输出或存储解析结果
    print(f"食材编码: {alim_code}, 成分编码: {const_code}, 含量: {teneur}")
    print(f"最小值状态: {min_missing}, 最大值状态: {max_missing}")
    print(f"可信度代码: {code_confiance}, 来源状态: {source_missing}\n")

3. 用lxml增强解析(容错+XPath支持)

如果原始XML格式混乱(比如未闭合标签、未转义字符),lxml的recover模式能极大提升容错性,还支持强大的XPath查询:

from lxml import etree
import codecs

# 开启容错解析模式
parser = etree.XMLParser(recover=True)

# 读取并解析XML
with codecs.open('your_xml_file.xml', 'r', 'windows-1252') as f:
    tree = etree.parse(f, parser)
    root = tree.getroot()

# 用XPath遍历所有COMPO节点
for compo in root.xpath('//COMPO'):
    # 用XPath提取文本(自动处理空列表)
    alim_code = compo.xpath('./alim_code/text()')[0].strip() if compo.xpath('./alim_code/text()') else None
    # 其他字段同理,这里简化示例
    print(f"食材编码: {alim_code}")

4. 核心注意事项

  • 编码匹配:XML指定了windows-1252编码,读取时必须用对应编码打开,否则会出现乱码。
  • 属性处理:像<min missing=" " />这种自闭合节点,没有文本内容,要通过读取missing属性来判断数据状态。
  • 特殊字符:提前修正XML中的非法字符是最优解,如果没法修改,再用容错解析库兜底。

内容的提问来源于stack exchange,提问作者DrCoolZic

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:05:29