XML转Pandas DataFrame报ParseError错误及通用转换方案咨询
回答
报错原因
XML官方规范明确要求所有标签属性的值必须用单引号或双引号包裹,你最初的XML文件中AUTHOR_UID = 1的属性值1没有加引号,属于不合法的XML格式,标准XML解析器(如你使用的xml.etree.ElementTree)会直接抛出格式校验错误,也就是你遇到的ParseError。
绕过方案
方案1:预处理XML补全属性引号
如果无法修改原始XML的生成逻辑,可以在解析前先读取XML文本,通过正则匹配给无引号的数字属性值自动补全引号,再用字符串解析XML即可,示例修改如下:
import re import pandas as pd import xml.etree.ElementTree as et def parse_XML(xml_file, df_cols): # 先读入文本做格式修复 with open(xml_file, "r", encoding="utf-8") as f: raw_xml = f.read() # 匹配无引号的数字属性值,补双引号 fixed_xml = re.sub(r'(\w+)\s*=\s*(\d+)', r'\1="\2"', raw_xml) xtree = et.ElementTree(et.fromstring(fixed_xml)) xroot = xtree.getroot() rows = [] for node in xroot: res = [] res.append(node.attrib.get(df_cols[0])) for el in df_cols[1:]: if node is not None and node.find(el) is not None: res.append(node.find(el).text) else: res.append(None) rows.append({df_cols[i]: res[i] for i, _ in enumerate(df_cols)}) out_df = pd.DataFrame(rows, columns=df_cols) return out_df
方案2:使用高容错性的XML解析器
改用BeautifulSoup搭配lxml解析器,它对不规范XML的容错性更高,无需预处理即可直接解析属性无引号的XML。
先安装依赖:pip install beautifulsoup4 lxml
修改后的解析代码示例:
from bs4 import BeautifulSoup import pandas as pd def parse_XML(xml_file, df_cols): with open(xml_file, "r", encoding="utf-8") as f: soup = BeautifulSoup(f, "lxml-xml") xroot = soup.dataset rows = [] for node in xroot.find_all("AUTHOR"): res = [] res.append(node.attrs.get(df_cols[0])) for el in df_cols[1:]: if node.find(el) is not None: res.append(node.find(el).text) else: res.append(None) rows.append({df_cols[i]: res[i] for i, _ in enumerate(df_cols)}) out_df = pd.DataFrame(rows, columns=df_cols) return out_df
无需硬编码的嵌套XML转DataFrame实现
核心思路是通过递归遍历所有节点,将嵌套的节点路径、节点属性自动扁平化为DataFrame的列名,不需要提前指定列名,只需要指定你要作为DataFrame行单位的节点标签即可,实现代码如下:
import pandas as pd from bs4 import BeautifulSoup def flatten_xml_node(node, parent_path="", sep="."): """递归扁平化XML节点为单层字典""" flat_dict = {} # 处理当前节点的所有属性 for attr_name, attr_val in node.attrs.items(): key = f"{parent_path}{sep}{attr_name}" if parent_path else attr_name flat_dict[key] = attr_val # 遍历子节点 for child in node.children: # 跳过纯文本空白节点 if child.name is None: continue child_path = f"{parent_path}{sep}{child.name}" if parent_path else child.name # 子节点没有后代节点时直接取文本内容 if len(child.find_all(recursive=False)) == 0: text_val = child.get_text(strip=True) flat_dict[child_path] = text_val if text_val else None # 有后代节点则递归处理 else: flat_dict.update(flatten_xml_node(child, child_path, sep)) return flat_dict def parse_nested_xml_to_df(xml_path, row_tag): """ 通用嵌套XML转DataFrame方法 :param xml_path: XML文件路径 :param row_tag: 作为DataFrame每一行对应节点的标签名,比如示例中的"AUTHOR" """ with open(xml_path, "r", encoding="utf-8") as f: soup = BeautifulSoup(f, "lxml-xml") # 提取所有行节点 row_nodes = soup.find_all(row_tag) # 扁平化所有节点生成数据集 records = [flatten_xml_node(node) for node in row_nodes] return pd.DataFrame(records) # 调用示例,不需要指定列名,只需要传入行节点标签即可 df = parse_nested_xml_to_df("your_xml_file.xml", row_tag="AUTHOR")
内容的提问来源于stack exchange,提问作者sree vatsav
相关产品推荐
相关产品推荐

