You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

XML转Pandas DataFrame报ParseError错误及通用转换方案咨询

回答

报错原因

XML官方规范明确要求所有标签属性的值必须用单引号或双引号包裹,你最初的XML文件中AUTHOR_UID = 1的属性值1没有加引号,属于不合法的XML格式,标准XML解析器(如你使用的xml.etree.ElementTree)会直接抛出格式校验错误,也就是你遇到的ParseError。

绕过方案

方案1:预处理XML补全属性引号

如果无法修改原始XML的生成逻辑,可以在解析前先读取XML文本,通过正则匹配给无引号的数字属性值自动补全引号,再用字符串解析XML即可,示例修改如下:

import re
import pandas as pd
import xml.etree.ElementTree as et

def parse_XML(xml_file, df_cols): 
    # 先读入文本做格式修复
    with open(xml_file, "r", encoding="utf-8") as f:
        raw_xml = f.read()
    # 匹配无引号的数字属性值,补双引号
    fixed_xml = re.sub(r'(\w+)\s*=\s*(\d+)', r'\1="\2"', raw_xml)
    xtree = et.ElementTree(et.fromstring(fixed_xml))
    xroot = xtree.getroot()
    rows = []
    
    for node in xroot: 
        res = []
        res.append(node.attrib.get(df_cols[0]))
        for el in df_cols[1:]: 
            if node is not None and node.find(el) is not None:
                res.append(node.find(el).text)
            else: 
                res.append(None)
        rows.append({df_cols[i]: res[i] 
                     for i, _ in enumerate(df_cols)})
    
    out_df = pd.DataFrame(rows, columns=df_cols)
        
    return out_df

方案2:使用高容错性的XML解析器

改用BeautifulSoup搭配lxml解析器,它对不规范XML的容错性更高,无需预处理即可直接解析属性无引号的XML。
先安装依赖:
pip install beautifulsoup4 lxml
修改后的解析代码示例:

from bs4 import BeautifulSoup
import pandas as pd

def parse_XML(xml_file, df_cols): 
    with open(xml_file, "r", encoding="utf-8") as f:
        soup = BeautifulSoup(f, "lxml-xml")
    xroot = soup.dataset
    rows = []
    
    for node in xroot.find_all("AUTHOR"): 
        res = []
        res.append(node.attrs.get(df_cols[0]))
        for el in df_cols[1:]: 
            if node.find(el) is not None:
                res.append(node.find(el).text)
            else: 
                res.append(None)
        rows.append({df_cols[i]: res[i] 
                     for i, _ in enumerate(df_cols)})
    
    out_df = pd.DataFrame(rows, columns=df_cols)
        
    return out_df

无需硬编码的嵌套XML转DataFrame实现

核心思路是通过递归遍历所有节点,将嵌套的节点路径、节点属性自动扁平化为DataFrame的列名,不需要提前指定列名,只需要指定你要作为DataFrame行单位的节点标签即可,实现代码如下:

import pandas as pd
from bs4 import BeautifulSoup

def flatten_xml_node(node, parent_path="", sep="."):
    """递归扁平化XML节点为单层字典"""
    flat_dict = {}
    # 处理当前节点的所有属性
    for attr_name, attr_val in node.attrs.items():
        key = f"{parent_path}{sep}{attr_name}" if parent_path else attr_name
        flat_dict[key] = attr_val
    # 遍历子节点
    for child in node.children:
        # 跳过纯文本空白节点
        if child.name is None:
            continue
        child_path = f"{parent_path}{sep}{child.name}" if parent_path else child.name
        # 子节点没有后代节点时直接取文本内容
        if len(child.find_all(recursive=False)) == 0:
            text_val = child.get_text(strip=True)
            flat_dict[child_path] = text_val if text_val else None
        # 有后代节点则递归处理
        else:
            flat_dict.update(flatten_xml_node(child, child_path, sep))
    return flat_dict

def parse_nested_xml_to_df(xml_path, row_tag):
    """
    通用嵌套XML转DataFrame方法
    :param xml_path: XML文件路径
    :param row_tag: 作为DataFrame每一行对应节点的标签名,比如示例中的"AUTHOR"
    """
    with open(xml_path, "r", encoding="utf-8") as f:
        soup = BeautifulSoup(f, "lxml-xml")
    # 提取所有行节点
    row_nodes = soup.find_all(row_tag)
    # 扁平化所有节点生成数据集
    records = [flatten_xml_node(node) for node in row_nodes]
    return pd.DataFrame(records)

# 调用示例,不需要指定列名,只需要传入行节点标签即可
df = parse_nested_xml_to_df("your_xml_file.xml", row_tag="AUTHOR")

内容的提问来源于stack exchange,提问作者sree vatsav

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 02:30:02