无需命名空间URI/URL,解析带前缀的XML文件(支持非Pandas方案)
问题
- 能否在不指定命名空间URI/URL的情况下使用
pandas.read_xml()? - 我的XML文件
studentinfo.xml中部分标签带有命名空间前缀(如stu:、st:),是否可以无需定义命名空间URI/URL,遍历并解析该文件的所有兄弟标签与子标签内容?也接受非Pandas的XML解析方案。
XML示例
<?xml version="1.0" encoding="UTF-8"?> <stu:StudentBreakdown> <stu:Studentdata> <stu:StudentScreening> <st:name>Sam Davies</st:name> <st:age>15</st:age> <st:hair>Black</st:hair> <st:eyes>Blue</st:eyes> <st:grade>10</st:grade> <st:teacher>Draco Malfoy</st:teacher> <st:dorm>Innovation Hall</st:dorm> </stu:StudentScreening> <stu:StudentScreening> <st:name>Cassie Stone</st:name> <st:age>14</st:age> <st:hair>Science</st:hair> <st:grade>9</st:grade> <st:teacher>Luna Lovegood</st:teacher> </stu:StudentScreening> <stu:StudentScreening> <st:name>Derek Brandon</st:name> <st:age>17</st:age> <st:eyes>green</st:eyes> <st:teacher>Ron Weasley</st:teacher> <st:dorm>Hogtie Manor</st:dorm> </stu:StudentScreening> </stu:Studentdata> </stu:StudentBreakdown>
现有代码
import pandas as pd from bs4 import BeautifulSoup with open('studentinfo.xml', 'r') as f: file = f.read() def parse_xml(file): soup = BeautifulSoup(file, 'xml') df1 = pd.DataFrame(columns=['StudentName', 'Age', 'Hair', 'Eyes', 'Grade', 'Teacher', 'Dorm']) all_items = soup.find_all('info') items_length = len(all_items) for index, info in enumerate(all_items): StudentName = info.find('<st:name>').text Age = info.find('<st:age>').text Hair = info.find('<st:hair>').text Eyes = info.find('<st:eyes>').text Grade = info.find('<st:grade>').text Teacher = info.find('<st:teacher>').text Dorm = info.find('<st:dorm>').text row = { 'StudentName': StudentName, 'Age': Age, 'Hair': Hair, 'Eyes': Eyes, 'Grade': Grade, 'Teacher': Teacher, 'Dorm': Dorm } df1 = df1.append(row, ingore_index=True) print(f'Appending row %s of %s' %(index+1, items_length)) return df1
期望输出
| Name | age | hair | eyes | grade | teacher | dorm | |
|---|---|---|---|---|---|---|---|
| 0 | Sam Davies | 15 | Black | Blue | 10 | Draco Malfoy | Innovation Hall |
| 1 | Cassie Stone | 14 | Science | N/A | 9 | Luna Lovegood | N/A |
| 2 | Derek Brandon | 17 | N/A | green | N/A | Ron Weasley | Hogtie Manor |
解决方案
一、Pandas方案(无需命名空间URI)
pandas.read_xml()可以通过XPath的local-name()函数忽略命名空间前缀,无需指定URI。核心是匹配标签的本地名称而非带前缀的完整标签名:
import pandas as pd # 读取XML并通过XPath忽略命名空间 df = pd.read_xml( 'studentinfo.xml', xpath='//*[local-name()="StudentScreening"]/*[local-name()="name" or local-name()="age" or local-name()="hair" or local-name()="eyes" or local-name()="grade" or local-name()="teacher" or local-name()="dorm"]', namespaces={} # 空命名空间字典,避免自动解析前缀 ).groupby(level=0).first().reset_index(drop=True) # 重命名列并填充缺失值 df.columns = ['Name', 'age', 'hair', 'eyes', 'grade', 'teacher', 'dorm'] df = df.fillna('N/A') print(df)
二、BeautifulSoup修正方案(无需命名空间URI)
你的现有代码存在标签查找错误,修正后可以直接匹配带前缀的标签,或提取标签的本地名称来遍历,无需命名空间URI:
import pandas as pd from bs4 import BeautifulSoup def parse_xml(file): soup = BeautifulSoup(file, 'xml') # 查找所有带stu:前缀的StudentScreening标签 all_items = soup.find_all('stu:StudentScreening') df1 = pd.DataFrame(columns=['Name', 'age', 'hair', 'eyes', 'grade', 'teacher', 'dorm']) for index, screening in enumerate(all_items): row = {} # 遍历子标签,提取冒号后的本地名称和文本 for tag in screening.find_all(): tag_name = tag.name.split(':')[-1] row[tag_name] = tag.text # 填充缺失字段为N/A for col in df1.columns: row[col] = row.get(col, 'N/A') # 追加行到DataFrame df1.loc[index] = row print(f'Appending row {index+1} of {len(all_items)}') return df1 # 读取文件并解析 with open('studentinfo.xml', 'r') as f: file_content = f.read() result_df = parse_xml(file_content) print(result_df)
这个方案会自动遍历每个学生节点下的所有子标签,不管是否带前缀,同时自动填充缺失字段,完全匹配你的期望输出。
内容的提问来源于stack exchange,提问作者PickleRick
相关产品推荐
相关产品推荐

