如何让Pandas to_xml()仅为根节点添加XML前缀而非所有行
解决Pandas to_xml仅为根节点添加命名空间前缀的问题
Pandas的to_xml()方法的prefix参数会给所有节点统一添加前缀,无法单独只给根节点设置。要实现仅根节点MT_InboundDeliveryDate带ns0前缀、其他节点无前缀的需求,可以通过以下两种方式实现:
方法1:字符串替换(快速实现)
先不设置prefix参数生成无前缀的XML,再手动替换根节点标签并添加命名空间声明:
import pandas as pd # 假设data为目标DataFrame namespaces = { 'ns0': "urn:sca:com:edi:mappings:aust:b2be:inbounddeliverydate" } # 生成无前缀的XML内容 xml_str = data.to_xml( index=False, root_name='MT_InboundDeliveryDate', row_name='Row' ) # 替换根节点标签,添加前缀与命名空间声明 namespace_decl = f'xmlns:ns0="{namespaces["ns0"]}"' xml_str = xml_str.replace( '<MT_InboundDeliveryDate>', f'<ns0:MT_InboundDeliveryDate {namespace_decl}>' ).replace( '</MT_InboundDeliveryDate>', '</ns0:MT_InboundDeliveryDate>' ) # 写入文件 with open('Inb.xml', 'w') as myfile: myfile.write(xml_str)
方法2:使用lxml库修改XML(更严谨)
如果XML结构复杂,字符串替换可能存在风险,推荐用lxml解析并修改XML:
import pandas as pd from lxml import etree namespaces = { 'ns0': "urn:sca:com:edi:mappings:aust:b2be:inbounddeliverydate" } # 生成无前缀的XML字节流 xml_bytes = data.to_xml( index=False, root_name='MT_InboundDeliveryDate', row_name='Row', encoding='utf-8' ) # 解析XML内容 root = etree.fromstring(xml_bytes) # 为根节点添加命名空间前缀 root.tag = f'{{{namespaces["ns0"]}}}{root.tag}' # 生成最终的XML字符串 xml_str = etree.tostring( root, encoding='utf-8', xml_declaration=True, pretty_print=True ).decode('utf-8') # 写入文件 with open('Inb.xml', 'w') as myfile: myfile.write(xml_str)
最终输出验证
两种方法生成的XML均符合需求:
<?xml version='1.0' encoding='utf-8'?> <ns0:MT_InboundDeliveryDate xmlns:ns0="urn:sca:com:edi:mappings:aust:b2be:inbounddeliverydate"> <Row> <InboundID>355555106537455</InboundID> <DocumentDate/> <LFDAT>19082022</LFDAT> </Row> <Row> <InboundID>35555552066774536</InboundID> <DocumentDate/> <LFDAT>03012023</LFDAT> </Row> </ns0:MT_InboundDeliveryDate>
内容的提问来源于stack exchange,提问作者Vatsal Patel
相关产品推荐
相关产品推荐

