You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas解析XML时多值字段出现NaN值的问题排查

问题:解析XML到DataFrame时多值字段出现NaN及行生成疑问

问题详情

原始XML文档

<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<MedicineResultOutput hits="1" offset="0" totalResults="90287">
    <SearchResults>
        <medicine id="1234" name="name1" lastModificationDate="2023-04-26T00:00:00Z" status="Launched">
            <CompanySource>gftr</CompanySource>
            <MainCompany>
                <Company>Com1</Company>
                <Company>Com2 Inc</Company>
                <Company>Com3 plc</Company>
            </MainCompany>
            <RegDes>
                <RegulatoryDesignation>RD1</RegulatoryDesignation>
                <RegulatoryDesignation>RD2</RegulatoryDesignation>
            </RegDes>
            <symptomPrimary>
                <symptom>symp1</symptom>
                <symptom>symp1</symptom>
                <symptom>symp3</symptom>
            </symptomPrimary>
            <ActPrim>
                <Action>actn1</Action>
                <Action>actn2</Action>
                <Action>actn3</Action>
            </ActPrim>
            <Tech>
                <Technology>tech1</Technology>
                <Technology>tech2</Technology>
            </Tech>
            <ThAreas>
                <TA>TS1</TA>
                <TA>TS2</TA>
            </ThAreas>
            <Summary>Some random summary</Summary>
            <ComSec>
                <Company>Comp23</Company>
                <Company>Comp34</Company>
            </ComSec>
            <symptomsSecondary>
                <symptom>SS2</symptom>
                <symptom>SS3</symptom>
            </symptomsSecondary>
            <ActionsSecondary>
                <Action>ACT223</Action>
                <Action>ACT567</Action>
            </ActionsSecondary>
            <AddedDate>1996-02-16T00:00:00Z</AddedDate>
        </medicine>
    </SearchResults>
</MedicineResultOutput>

期望的DataFrame输出

id    name    lastModificationDate status   CompanySource  MainCompany               RegDes    symptomPrimary        ActPrim             Tech           ThAreas     ComSec           symptomsSecondary  ActionsSecondary   AddedDate
1234  name1   2023-04-26T00:00:00Z Launched  gftr     [Com1,Com2 Inc,Com3 plc]  [RD1,RD2]  [symp1,symp1,symp3]  [actn1,actn2,actn3]  [tech1,tech2]  [TS1,TS2]  [Comp23,Comp34]   [SS2,SS3]          [ACT223,ACT567]     1996-02-16T00:00:00Z

当前代码及问题

使用pandas的read_xml解析:

df = pd.read_xml(response.text, xpath='.//medicine')

遇到的问题:

  • 所有列都显示,但MainCompany、symptomPrimary等包含多值的字段出现NaN值
  • 原本以为多值字段会生成多行数据(比如MainCompany有3个值就生成3行),但实际没有

解决方案

1. 多值字段出现NaN的原因

pd.read_xml默认只会解析直接子节点的文本内容,对于MainCompany这种包含子节点(<Company>)的父节点,它无法自动提取子节点的集合,所以返回NaN。要提取这类嵌套结构的多值字段,需要自定义解析逻辑。

2. 实现目标解析的代码

结合xml.etree.ElementTree手动解析XML,提取每个medicine节点的属性和嵌套子节点的列表:

import pandas as pd
import xml.etree.ElementTree as ET

# 解析XML内容
root = ET.fromstring(response.text)
medicines = []

for med in root.findall('.//medicine'):
    # 提取medicine节点的属性
    med_data = {
        'id': med.get('id'),
        'name': med.get('name'),
        'lastModificationDate': med.get('lastModificationDate'),
        'status': med.get('status'),
        'CompanySource': med.find('CompanySource').text if med.find('CompanySource') is not None else None,
        'Summary': med.find('Summary').text if med.find('Summary') is not None else None,
        'AddedDate': med.find('AddedDate').text if med.find('AddedDate') is not None else None
    }
    
    # 提取嵌套的多值字段
    main_companies = [comp.text for comp in med.find('MainCompany').findall('Company')] if med.find('MainCompany') is not None else []
    med_data['MainCompany'] = main_companies
    
    reg_des = [rd.text for rd in med.find('RegDes').findall('RegulatoryDesignation')] if med.find('RegDes') is not None else []
    med_data['RegDes'] = reg_des
    
    primary_symptoms = [s.text for s in med.find('symptomPrimary').findall('symptom')] if med.find('symptomPrimary') is not None else []
    med_data['symptomPrimary'] = primary_symptoms
    
    primary_actions = [a.text for a in med.find('ActPrim').findall('Action')] if med.find('ActPrim') is not None else []
    med_data['ActPrim'] = primary_actions
    
    techs = [t.text for t in med.find('Tech').findall('Technology')] if med.find('Tech') is not None else []
    med_data['Tech'] = techs
    
    th_areas = [ta.text for ta in med.find('ThAreas').findall('TA')] if med.find('ThAreas') is not None else []
    med_data['ThAreas'] = th_areas
    
    sec_companies = [comp.text for comp in med.find('ComSec').findall('Company')] if med.find('ComSec') is not None else []
    med_data['ComSec'] = sec_companies
    
    secondary_symptoms = [s.text for s in med.find('symptomsSecondary').findall('symptom')] if med.find('symptomsSecondary') is not None else []
    med_data['symptomsSecondary'] = secondary_symptoms
    
    secondary_actions = [a.text for a in med.find('ActionsSecondary').findall('Action')] if med.find('ActionsSecondary') is not None else []
    med_data['ActionsSecondary'] = secondary_actions
    
    medicines.append(med_data)

# 转换为DataFrame
df = pd.DataFrame(medicines)
print(df)

3. 关于多行生成的处理

pd.read_xml默认不会根据多值字段拆分生成多行,若需要将多值字段拆分成单行对应单个值的格式,可使用explode方法:

# 根据MainCompany字段拆分多行
df_exploded = df.explode('MainCompany', ignore_index=True)
print(df_exploded)

执行后会生成3行数据,每行对应一个MainCompany的值,其他字段保持不变。


内容的提问来源于stack exchange,提问作者Dcook

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 06:05:15