墨西哥电子发票XML转CSV的Python实现问题排查
解决墨西哥CFDI XML转CSV的Python常见问题
问题1:写入CSV的是DOM元素对象而非属性值
直接写入元素对象会导致CSV出现<Element 'xxx' at 0x...>这类无效内容,必须明确提取属性值:
- 用元素的
attrib字典获取属性,比如element.attrib['属性名'] - 可选属性用
get方法兜底,避免KeyError:element.attrib.get('属性名', '')
问题2:无法提取Total和Fecha属性
这类属性属于根节点cfdi:Comprobante,核心问题是未处理XML命名空间。CFDI标准命名空间为http://www.sat.gob.mx/cfd/3,查找节点时必须指定:
- 注册命名空间前缀,或直接用完整URI查找节点
- 示例:
tree.find('.//{http://www.sat.gob.mx/cfd/3}Comprobante')
问题3:多cfdi:Concepto时列数不匹配
发票含多个商品/服务条目时,需将主发票信息与每个Concepto条目合并为一行,而非把所有Concepto信息塞进同一行:
- 遍历每个
cfdi:Concepto节点 - 每次循环将主发票字段(Fecha、Total等)与当前Concepto字段(Descripcion、Cantidad等)拼接成一行数据
完整可运行代码示例
import csv import xml.etree.ElementTree as ET # 注册CFDI命名空间,解决前缀识别问题 ET.register_namespace('cfdi', 'http://www.sat.gob.mx/cfd/3') NS = {'cfdi': 'http://www.sat.gob.mx/cfd/3'} def cfdi_to_csv(xml_path, csv_path): # 读取XML文件 tree = ET.parse(xml_path) root = tree.getroot() # 获取主发票核心节点 comprobante = root.find('.//cfdi:Comprobante', NS) if not comprobante: raise ValueError("未找到cfdi:Comprobante节点") # 提取主发票固定字段 invoice_base = { 'Fecha': comprobante.attrib.get('Fecha', ''), 'Total': comprobante.attrib.get('Total', ''), 'Serie': comprobante.attrib.get('Serie', ''), 'Folio': comprobante.attrib.get('Folio', '') } # 获取所有商品条目节点 conceptos = root.findall('.//cfdi:Concepto', NS) if not conceptos: raise ValueError("未找到cfdi:Concepto节点") # 定义CSV表头,主字段+商品字段 headers = list(invoice_base.keys()) + ['Descripcion', 'Cantidad', 'Importe', 'ClaveProdServ'] # 写入CSV文件 with open(csv_path, 'w', newline='', encoding='utf-8') as csvfile: writer = csv.DictWriter(csvfile, fieldnames=headers) writer.writeheader() # 遍历每个商品条目,生成一行数据 for concepto in conceptos: row = invoice_base.copy() row.update({ 'Descripcion': concepto.attrib.get('Descripcion', ''), 'Cantidad': concepto.attrib.get('Cantidad', ''), 'Importe': concepto.attrib.get('Importe', ''), 'ClaveProdServ': concepto.attrib.get('ClaveProdServ', '') }) writer.writerow(row) # 使用示例 cfdi_to_csv('factura.xml', 'factura.csv')
关键说明
- 命名空间处理:通过
register_namespace和NS字典,确保正确识别带cfdi:前缀的节点 - 鲁棒性保障:全程用
attrib.get()处理属性,避免因属性缺失导致程序崩溃 - 多条目兼容:每个Concepto对应一行CSV,主发票信息重复填充,彻底解决列数不匹配问题
- 格式兼容性:指定
utf-8编码避免乱码,newline=''保证CSV在Windows/Linux下格式一致
内容的提问来源于stack exchange,提问作者rwffh
相关产品推荐
相关产品推荐

