You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无需命名空间URI/URL,解析带前缀的XML文件(支持非Pandas方案)

问题
  1. 能否在不指定命名空间URI/URL的情况下使用pandas.read_xml()?
  2. 我的XML文件studentinfo.xml中部分标签带有命名空间前缀(如stu:、st:),是否可以无需定义命名空间URI/URL,遍历并解析该文件的所有兄弟标签与子标签内容?也接受非Pandas的XML解析方案。

XML示例

<?xml version="1.0" encoding="UTF-8"?>
<stu:StudentBreakdown>
<stu:Studentdata>
    <stu:StudentScreening>
        <st:name>Sam Davies</st:name>
        <st:age>15</st:age>
        <st:hair>Black</st:hair>
        <st:eyes>Blue</st:eyes>
        <st:grade>10</st:grade>
        <st:teacher>Draco Malfoy</st:teacher>
        <st:dorm>Innovation Hall</st:dorm>
    </stu:StudentScreening>
    <stu:StudentScreening>
        <st:name>Cassie Stone</st:name>
        <st:age>14</st:age>
        <st:hair>Science</st:hair>
        <st:grade>9</st:grade>
        <st:teacher>Luna Lovegood</st:teacher>
    </stu:StudentScreening>
    <stu:StudentScreening>
        <st:name>Derek Brandon</st:name>
        <st:age>17</st:age>
        <st:eyes>green</st:eyes>
        <st:teacher>Ron Weasley</st:teacher>
        <st:dorm>Hogtie Manor</st:dorm>
    </stu:StudentScreening>
</stu:Studentdata>
</stu:StudentBreakdown>

现有代码

import pandas as pd
from bs4 import BeautifulSoup
with open('studentinfo.xml', 'r') as f:
    file = f.read()  

def parse_xml(file):
    soup = BeautifulSoup(file, 'xml')
    df1 = pd.DataFrame(columns=['StudentName', 'Age', 'Hair', 'Eyes', 'Grade', 'Teacher', 'Dorm'])
    all_items = soup.find_all('info')
    items_length = len(all_items)
    for index, info in enumerate(all_items):
        StudentName = info.find('<st:name>').text
        Age = info.find('<st:age>').text
        Hair = info.find('<st:hair>').text
        Eyes = info.find('<st:eyes>').text
        Grade = info.find('<st:grade>').text
        Teacher = info.find('<st:teacher>').text
        Dorm = info.find('<st:dorm>').text
      row = {
            'StudentName': StudentName,
            'Age': Age,
            'Hair': Hair,
            'Eyes': Eyes,
            'Grade': Grade,
            'Teacher': Teacher,
            'Dorm': Dorm
        }
        
        df1 = df1.append(row, ingore_index=True)
        print(f'Appending row %s of %s' %(index+1, items_length))
    
    return df1  

期望输出

Nameagehaireyesgradeteacherdorm
0Sam Davies15BlackBlue10Draco MalfoyInnovation Hall
1Cassie Stone14ScienceN/A9Luna LovegoodN/A
2Derek Brandon17N/AgreenN/ARon WeasleyHogtie Manor

解决方案

一、Pandas方案(无需命名空间URI)

pandas.read_xml()可以通过XPath的local-name()函数忽略命名空间前缀,无需指定URI。核心是匹配标签的本地名称而非带前缀的完整标签名:

import pandas as pd

# 读取XML并通过XPath忽略命名空间
df = pd.read_xml(
    'studentinfo.xml',
    xpath='//*[local-name()="StudentScreening"]/*[local-name()="name" or local-name()="age" or local-name()="hair" or local-name()="eyes" or local-name()="grade" or local-name()="teacher" or local-name()="dorm"]',
    namespaces={}  # 空命名空间字典,避免自动解析前缀
).groupby(level=0).first().reset_index(drop=True)

# 重命名列并填充缺失值
df.columns = ['Name', 'age', 'hair', 'eyes', 'grade', 'teacher', 'dorm']
df = df.fillna('N/A')

print(df)

二、BeautifulSoup修正方案(无需命名空间URI)

你的现有代码存在标签查找错误,修正后可以直接匹配带前缀的标签,或提取标签的本地名称来遍历,无需命名空间URI:

import pandas as pd
from bs4 import BeautifulSoup

def parse_xml(file):
    soup = BeautifulSoup(file, 'xml')
    # 查找所有带stu:前缀的StudentScreening标签
    all_items = soup.find_all('stu:StudentScreening')
    df1 = pd.DataFrame(columns=['Name', 'age', 'hair', 'eyes', 'grade', 'teacher', 'dorm'])
    
    for index, screening in enumerate(all_items):
        row = {}
        # 遍历子标签,提取冒号后的本地名称和文本
        for tag in screening.find_all():
            tag_name = tag.name.split(':')[-1]
            row[tag_name] = tag.text
        # 填充缺失字段为N/A
        for col in df1.columns:
            row[col] = row.get(col, 'N/A')
        # 追加行到DataFrame
        df1.loc[index] = row
        print(f'Appending row {index+1} of {len(all_items)}')
    
    return df1

# 读取文件并解析
with open('studentinfo.xml', 'r') as f:
    file_content = f.read()

result_df = parse_xml(file_content)
print(result_df)

这个方案会自动遍历每个学生节点下的所有子标签,不管是否带前缀,同时自动填充缺失字段,完全匹配你的期望输出。


内容的提问来源于stack exchange,提问作者PickleRick

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 01:16:10