如何读取XML文件并生成指定格式的Pandas DataFrame
问题原因
你现有代码存在两个核心错误:
- 导入的XML解析模块别名为
ET,后续调用时误写为et,会直接报命名错误 row.get()方法用于获取XML节点的属性值,但你提供的XML中Application_ID、Product_Type、Product_ID都是ROW节点的子节点,值存储在子节点的文本中,无法用get方法读取
修正后可直接运行的代码
import xml.etree.cElementTree as ET import pandas as pd path = 'product.xml' # 定义需要提取的字段顺序 extract_cols = ['Application_ID', 'Product_Type', 'Product_ID'] # 解析XML文件 root = ET.parse(path).getroot() row_nodes = root.findall('.//ROW') # 遍历提取子节点文本 xml_content = [] for row in row_nodes: current_row = [] for col in extract_cols: target_node = row.find(col) current_row.append(target_node.text if target_node is not None else '') xml_content.append(current_row) # 生成DataFrame并重命名列 df = pd.DataFrame(xml_content, columns=extract_cols) df.columns = df.columns.str.replace('_', ' ') # 输出无行索引的表格,匹配你需要的格式 print(df.to_string(index=False))
输出效果
运行上述代码后会直接输出你要求的格式:
Application ID Product Type Product ID BBC#:1010 1 32 NBA#:1111 2 22 BBC#:1212 1 63 NBA#:2210 2 22
内容的提问来源于stack exchange,提问作者Romeo Botanpi
相关产品推荐
相关产品推荐

